System

The system addresses the challenge of real-time feedback in online negotiations by analyzing voice and video data to enhance communication effectiveness through emotion icons and speaking time ratios, improving negotiation quality.

JP2026017326APending Publication Date: 2026-02-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118108
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-02-04

AI Technical Summary

Technical Problem

Traditional online business negotiations and interviews lack real-time feedback on reactions and emotions, making it difficult to maintain effective communication and consistency in topics, which hinders performance.

Method used

A system that acquires and analyzes voice and video data in real-time using a voice recognition engine and facial expression analysis engine, displaying emotion icons, calculating speaking time ratios, and providing evaluation scores and suggestions to enhance communication effectiveness.

Benefits of technology

Enables users to grasp the other party's reactions and emotions in real-time, improving the quality of online business negotiations and interviews through enhanced communication support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017326000001_ABST
    Figure 2026017326000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring voice data and video data; means for passing the acquired voice data to a voice recognition engine and specifying utterance content and an utterer; means for passing the acquired video data to a facial expression analysis engine and specifying an emotion from a facial expression; means for totaling an utterance time for each utterer and calculating a ratio of a speaking time; means for displaying a text conversation log in real time; means for displaying an emotion analysis result as an emotion icon; and means for calculating and displaying an evaluation score based on the conversation log and emotion analysis data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Unlike face-to-face meetings, traditional online business negotiations and interviews make it difficult to control what and when to talk, making it difficult to properly read the other person's reactions. Furthermore, missing questions and maintaining consistency in topics can be challenging. This can prevent many business people from performing at their best. The objective of this invention is to provide a system that supports effective communication by providing users with the necessary information in real time, in order to improve the quality of their online business negotiations and interviews. [Means for solving the problem]

[0005] The present invention provides a system including means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine to identify the content of speech and the speaker, means for passing the acquired video data to a facial expression analysis engine to identify emotions from facial expressions, means for aggregating the speaking time of each speaker and calculating the percentage of speaking time, means for displaying a text-converted conversation log in real time, means for displaying the emotion analysis results as emotion icons, means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data, means for analyzing the progress of the conversation and generating and displaying suggestions for what to say next, means for accessing saved data and replaying and analyzing past business negotiations or interviews, and means for searching for and displaying the meaning of a selected word from a dictionary database, thereby enabling users to obtain necessary information in real time and communicate more effectively.

[0006] "Audio data" refers to sound information captured through a microphone during online business negotiations or interviews.

[0007] "Video data" refers to video information captured through a camera during online business negotiations or interviews.

[0008] A "voice recognition engine" is a program or algorithm that analyzes voice data and converts it into text.

[0009] "Speech content" is text data of words spoken in a conversation that is analyzed by a voice recognition engine.

[0010] The "speaker" is the person who actually made the statement, and is the target identified by the voice recognition engine.

[0011] An "facial expression analysis engine" is a program or algorithm that analyzes video data and identifies emotions from individual facial expressions.

[0012] "Emotion" is a psychological state that can be identified by a facial expression analysis engine or a tone of voice analysis engine.

[0013] "Speech time" refers to the amount of time a speaker speaks aloud during a conversation.

[0014] A "textualized conversation log" is a record of the conversation content converted into text format by a voice recognition engine.

[0015] An "emotion icon" is an icon that visually represents the results of emotion analysis and indicates the type of emotion to the user.

[0016] The "evaluation score" is a score based on conversation logs and sentiment analysis data, and is a numerical representation of a user's performance.

[0017] "Suggestions" are recommendations for what the user should say or do next.

[0018] A "dictionary database" is a database in which the meanings and definitions of words and phrases can be searched.

[0019] "Replaying a business meeting" refers to an operation in which a user plays back and checks saved meeting data. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0022] First, the terms used in the following description will be explained.

[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0028] [First embodiment]

[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0041] This invention relates to a system that supports online business negotiations and online interviews by acquiring and analyzing voice and video data in real time. This system is mainly composed of the following elements: a means for acquiring voice and video data, a voice recognition engine, a facial expression analysis engine, a means for displaying emotional icons, a means for automatically converting conversation logs into text, a means for calculating speech time and evaluation scores, a suggestion function, and a means for accessing a dictionary database.

[0042] Explanation of program processing

[0043] Setup and Connection

[0044] The server connects the user's device to a major web meeting product (such as a general web conferencing system). When the user starts a web meeting, the device connects to the server and launches a support tool, which starts collecting data for business negotiations or interviews.

[0045] Acquiring audio and video data

[0046] The device captures audio and video data in real time from the microphone and camera. This data is then sent directly to the server in streaming format. The user is then able to proceed with a normal web meeting without being aware of anything.

[0047] Speech recognition and talking time percentage calculation

[0048] The server passes the voice data to a voice recognition engine, which identifies the content of the speech and the speaker. The voice recognition engine analyzes the data, generates text data, and records the content of the speech and speaker information. The server then aggregates the speaking time for each speaker and calculates the "talking time ratio" for the user and the other party. This information is sent to the device in real time and displayed as an overlay on the screen.

[0049] sentiment analysis

[0050] The server passes the video data to a facial expression analysis engine, which identifies emotions from facial expressions. The audio data is also passed to a tone of voice analysis engine, which identifies emotions from the tone of voice. These emotion data are integrated and displayed on the user's screen as an emotion icon.

[0051] Automatic text conversion of conversation logs

[0052] The server records the text data obtained from the speech recognition engine in real time as a conversation log, which is periodically sent to the terminal and displayed on the user's screen in real time.

[0053] Calculation of evaluation score

[0054] The server evaluates the conversation logs and sentiment analysis data based on best practice rules to calculate a rating score, which is displayed in real time on the user's device and suggests areas for improvement if necessary.

[0055] Quick Dictionary Function

[0056] When a user selects a word they don't understand, the device sends it to the server, which searches for the word's meaning in the dictionary database and sends the results to the user's device, where the user can see the meaning of the selected word in a pop-up window.

[0057] Suggestion function

[0058] The server analyzes the conversation and its progress, and generates suggestions for what to say next and what questions to ask. These suggestions are sent to the user's device in real time and displayed in a pop-up window.

[0059] Retrospective training function

[0060] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and practice repeatedly using simulation videos that include areas of failure and areas for improvement.

[0061] Specific examples

[0062] For example, consider a case where a user has a 10-minute online business meeting. The device sends the audio and video data of the meeting to the server in real time. The server analyzes the audio data and obtains the results: the user spoke for 6 minutes and the other party spoke for 4 minutes. Based on this, the speaking time ratios of 68% vs. 32% are displayed on the screen. At the same time, the content of the conversation is converted into text and displayed in real time on the user's screen as a conversation log. The server also analyzes the other party's facial expressions and tone of voice, and displays emotion icons such as "skepticism" or "interest."

[0063] In this way, the present invention is a multifunctional, real-time support system that enables users to conduct online business negotiations and online interviews more effectively.

[0064] The processing flow will be explained below.

[0065] Explanation of program processing

[0066] Setup and Connection

[0067] Step 1:

[0068] The server connects users' devices and major web meeting products.

[0069] Step 2:

[0070] When a user starts a Web meeting, the terminal notifies the server of the start of the session.

[0071] Step 3:

[0072] The server transmits the initial setting data of the tool to the user's terminal.

[0073] Acquiring audio and video data

[0074] Step 1:

[0075] The device captures audio and video data in real time from the microphone and camera.

[0076] Step 2:

[0077] The terminal transmits the acquired audio and video data to the server in streaming format.

[0078] Speech recognition and talking time percentage calculation

[0079] Step 1:

[0080] The server passes the received voice data to the voice recognition engine.

[0081] Step 2:

[0082] The voice recognition engine analyzes the voice data and identifies what is being said and who is speaking.

[0083] Step 3:

[0084] The server records the analyzed text data and speaker information.

[0085] Step 4:

[0086] The server aggregates the speaking time for each speaker and calculates the "proportion of speaking time" for the user and the other party.

[0087] Step 5:

[0088] The device overlays the percentage of speaking time received from the server on the screen.

[0089] sentiment analysis

[0090] Step 1:

[0091] The server passes the video data to the facial expression analysis engine.

[0092] Step 2:

[0093] The facial expression analysis engine analyzes video data and identifies emotions from facial expressions.

[0094] Step 3:

[0095] The server passes the audio data to a tone of voice analysis engine.

[0096] Step 4:

[0097] A voice tone analysis engine analyzes audio data and identifies emotions from the tone of voice.

[0098] Step 5:

[0099] The server integrates the obtained emotion data and generates an emotion icon.

[0100] Step 6:

[0101] The device overlays the emotion icons received from the server on the screen.

[0102] Automatic text conversion of conversation logs

[0103] Step 1:

[0104] The server records the conversation content obtained from the voice recognition engine as a text log in real time.

[0105] Step 2:

[0106] The server periodically sends the text log to the user's terminal.

[0107] Step 3:

[0108] The device overlays the received conversation log on the screen in real time.

[0109] Calculation of evaluation score

[0110] Step 1:

[0111] The server analyzes the conversation logs and sentiment analysis data and evaluates them based on best practice rules.

[0112] Step 2:

[0113] The server calculates the evaluation score and sends the score to the user's terminal.

[0114] Step 3:

[0115] The device displays the received evaluation score on the screen and suggests areas for improvement if necessary.

[0116] Quick Dictionary Function

[0117] Step 1:

[0118] When the user selects a word they do not understand, the terminal sends the word to the server.

[0119] Step 2:

[0120] The server searches the dictionary database for the meaning of the word and sends the result to the terminal.

[0121] Step 3:

[0122] The device will pop up the meaning of the word on the screen.

[0123] Suggestion function

[0124] Step 1:

[0125] The server analyzes the conversation content and progress and generates suggestions for what to say next.

[0126] Step 2:

[0127] The server sends the generated suggestions to the user's terminal.

[0128] Step 3:

[0129] The device displays the received suggestions in a pop-up format on the screen.

[0130] Retrospective training function

[0131] Step 1:

[0132] The server stores all data from the meeting (audio, video, conversation logs, sentiment analysis).

[0133] Step 2:

[0134] The server provides a review interface that users can access after the meeting has ended.

[0135] Step 3:

[0136] Users can access the saved data to replay and analyze past business meetings and interviews.

[0137] Step 4:

[0138] Users can repeatedly practice the areas where they failed or the points they want to improve using simulation videos.

[0139] Example 1

[0140] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0141] In recent years, remote work and online business negotiations and interviews have become more common, but communication in these online environments is considered more difficult than face-to-face communication. It is particularly difficult to grasp the other person's reactions and emotions in real time, making it difficult to effectively progress and evaluate the dialogue. Therefore, there is a demand for a system that supports online business negotiations and interviews, with a function that can analyze the other person's reactions and emotions in real time and enable users to communicate effectively.

[0142] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0143] In this invention, the server includes means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine and identifying the content of the speech and the speaker, means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions, means for aggregating the speaking time for each speaker and calculating the percentage of speaking time, means for displaying a textualized conversation log in real time, means for displaying the emotion analysis results as emotion icons, means for recording the text data acquired from the voice recognition engine as a conversation log and transmitting it to a terminal, and means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data. This makes it possible to grasp the other party's reactions and emotions in real time and provide an environment in which users can communicate effectively online.

[0144] "Audio data" refers to sound information acquired through a microphone during online business negotiations or online interviews.

[0145] "Video data" refers to video information captured through a camera during online business negotiations or interviews.

[0146] A "voice recognition engine" is a general term for technology that automatically identifies the content and speaker of an utterance from acquired voice data and converts it into text data.

[0147] "Facial expression analysis engine" is a general term for technology that analyzes facial expressions from acquired video data and identifies emotions from them.

[0148] "Speech duration" refers to the length of time during which a particular speaker speaks.

[0149] "Text conversion" is the process of analyzing audio data and recording what is being said as written information.

[0150] A "conversation log" is a record of text data generated by a voice recognition engine, recorded in chronological order.

[0151] "Emotion icons" are shapes or facial expression icons that visually represent the results of the emotion analysis engine.

[0152] The "evaluation score" is a numerical evaluation value that represents the quality and effectiveness of a conversation based on conversation logs and sentiment analysis data.

[0153] "Suggestion" is a function that analyzes the content and progress of a conversation and suggests what to talk about next or what questions to ask.

[0154] A "dictionary database" is a database that stores the meanings and explanations of words.

[0155] "Real-time display" refers to the act of displaying acquired data and analysis results instantly with almost no delay.

[0156] "Cloud storage" is an online storage service for storing data over the Internet.

[0157] This invention is a system that supports online business negotiations and online interviews by capturing and analyzing audio and video data in real time. This system has the following configuration, including the following main hardware and software elements: a microphone, a camera, a voice recognition engine, a facial expression analysis engine, a suggestion function, and a dictionary database access means.

[0158] Setup and Connection

[0159] The server connects the user's device to the main web meeting product. When the user starts a web meeting, the device connects to the server and launches the support tool, which starts collecting data for the business meeting or interview.

[0160] Acquiring audio and video data

[0161] The device uses a microphone and camera to capture audio and video data in real time, which is then sent in streaming format to a server for analysis.

[0162] Speech recognition and talking time percentage calculation

[0163] The server passes the voice data to a voice recognition engine, which identifies the content of the speech and the speaker. The voice recognition engine converts the voice data into text data and records the content of the speech and speaker information. The server then aggregates the speaking time for each speaker and calculates the "talking time ratio" for the user and the other party. This information is sent to the device in real time and displayed as an overlay on the screen.

[0164] As a specific example of use, the voice recognition engine uses the Google Cloud Speech-to-Text API to convert actual voice data into text with extremely high accuracy.

[0165] sentiment analysis

[0166] The server passes the video data to a facial expression analysis engine, which identifies emotions from facial expressions. The audio data is passed to a tone of voice analysis engine, which identifies emotions from the tone of voice. These emotion data are integrated and displayed on the user's screen as an emotion icon.

[0167] As a specific example of use, the facial expression analysis engine uses Microsoft Azure Face API, and the voice tone analysis engine uses IBM Watson Tone Analyzer.

[0168] Automatic text conversion of conversation logs

[0169] The server records the text data obtained from the speech recognition engine in real time as a conversation log, which is periodically sent to the terminal and displayed on the user's screen in real time.

[0170] Calculation of evaluation score

[0171] The server evaluates the conversation log and sentiment analysis data using best practice rules to calculate a score, which is displayed on the device in real time and suggests areas for improvement if necessary.

[0172] Quick Dictionary Function

[0173] When a user selects a word they do not understand, the device sends it to the server, which searches for the meaning of the word in a dictionary database and sends the results to the device, allowing the user to check the meaning of the selected word in a pop-up window.

[0174] Suggestion function

[0175] The server analyzes the conversation and progress, and generates suggestions for what to say next and what questions to ask. These suggestions are sent to the device in real time and displayed in a pop-up format.

[0176] Retrospective training function

[0177] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and practice repeatedly using simulation videos that include areas of failure and areas for improvement.

[0178] Specific examples

[0179] For example, if a user has a 10-minute online business meeting, the device sends the audio and video data of the meeting to the server in real time. The server analyzes the audio data and determines that the user spoke for 6 minutes and the other party spoke for 4 minutes. Based on this, the speaking time ratio of 68% vs. 32% is displayed on the screen. At the same time, the content of the conversation is converted into text and displayed on the user's screen in real time as a conversation log. The server also analyzes the other party's facial expressions and tone of voice, and displays emotion icons such as "suspicious" and "interested."

[0180] Prompt Sentence Examples

[0181] "Please explain about a system that allows users to check emotional icons and conversation logs in real time during online business negotiations and suggests ways to improve their conversation skills based on the other party's reactions."

[0182] In this way, the present invention is a multi-functional, real-time support system that enables users to conduct online business negotiations and online interviews more effectively.

[0183] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0184] Step 1: Setup and Connection

[0185] The server connects the user's device to the main Web meeting product. When the user starts a Web meeting, the device works with the server to launch the support tool.

[0186] Specific behavior:

[0187] Input: User action to start a web meeting

[0188] Data processing: The server connects to the meeting using the API of the web conferencing system.

[0189] Output: Gets meeting session information and sends it to the device

[0190] Step 2: Acquiring audio and video data

[0191] The device uses a microphone and camera to capture audio and video data in real time, and transmits this data to a server in streaming format.

[0192] Specific behavior:

[0193] Input: Audio data (from microphone), video data (from camera)

[0194] Data Processing: Real-time acquisition and synchronization of audio and video data

[0195] Output: Sends audio and video data to the server in streaming format

[0196] Step 3: Speech recognition and speaking time calculation

[0197] The server passes the voice data to a speech recognition engine to identify the content and speaker, then aggregates the speaking time for each speaker and calculates the "speaking time percentage."

[0198] Specific behavior:

[0199] Input: A stream of audio data

[0200] Data calculation: Converting voice data into text using the Google Cloud Speech-to-Text API, and identifying the content and speaker

[0201] Output: Send the aggregated results to the terminal and overlay them on the screen

[0202] Step 4: Sentiment analysis

[0203] The server passes the video data to an expression analysis engine to determine emotions from facial expressions, and passes the audio data to a tone of voice analysis engine to determine emotions.

[0204] Specific behavior:

[0205] Input: Video data (to facial expression analysis engine), audio data (to voice tone analysis engine)

[0206] Data calculation: Facial expression analysis using Microsoft Azure Face API, voice tone analysis using IBM Watson Tone Analyzer

[0207] Output: Generate emoticons and send them to the device

[0208] Step 5: Automatic transcription of conversation logs

[0209] The server records the text data obtained from the voice recognition engine as a conversation log and periodically sends it to the terminal.

[0210] Specific behavior:

[0211] Input: Text data from the speech recognition engine

[0212] Data processing: Recorded in a database as a conversation log

[0213] Output: Real-time updated conversation log sent to device

[0214] Step 6: Calculating the evaluation score

[0215] The server calculates a rating score using best practice rules based on conversation logs and sentiment analysis data.

[0216] Specific behavior:

[0217] Input: Conversation logs, sentiment analysis data

[0218] Data calculations: evaluated based on best practice rules

[0219] Output: Send the evaluation score to the terminal and display it.

[0220] Step 7: Quick Dictionary Function

[0221] The terminal sends the word selected by the user to the server, which searches for the meaning of the word from a dictionary database and returns it.

[0222] Specific behavior:

[0223] Input: Selected word (from terminal)

[0224] Data operations: Retrieving word meanings from dictionary databases (e.g., Merriam-Webster API)

[0225] Output: Send word meanings to terminal and display

[0226] Step 8: Suggestions

[0227] Based on the analysis of the conversation content and the progress, the server generates the next topic to talk about and questions, and sends them to the user's device in real time.

[0228] Specific behavior:

[0229] Input: Conversation content and progress data

[0230] Data Computation: Using generative AI models to generate suggestions

[0231] Output: Send suggestions to the device and display them in a popup

[0232] Step 9: Retrospective Training Function

[0233] The server stores all data from the meeting and allows users to access a review interface after the meeting, allowing them to replay and analyze past data and identify areas for improvement.

[0234] Specific behavior:

[0235] Input: Meeting audio, video, conversation logs, and sentiment analysis data

[0236] Data processing: Save data to cloud storage

[0237] Output: Provides a retrospective interface that allows users to play back and analyze data

[0238] (Application example 1)

[0239] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0240] Conventional customer service systems lack real-time feedback and support to enable store clerks to communicate smoothly with customers. This can lead to an inability to respond quickly to customer needs, which can lead to a decline in customer satisfaction. This can also have a negative impact on store sales and repeat customer rates. The present invention aims to solve these problems and provide a system that enables store clerks to communicate more effectively and smoothly with customers.

[0241] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0242] In this invention, the server includes means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine and identifying the content of the utterance and the speaker, means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions, means for aggregating the speaking time for each speaker and calculating the percentage of speaking time, means for displaying a text-converted conversation log in real time, means for displaying the emotion analysis results as an emotion icon, means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data, and means for analyzing customer emotions and conversation content during customer service in real time and generating and displaying suggestions for what to say next. This enables store clerks to analyze communication with customers in real time and appropriately determine the optimal response and next action, thereby improving customer satisfaction.

[0243] "Audio data" refers to audio information recorded in digital format.

[0244] "Video data" refers to video information recorded in digital format.

[0245] A "voice recognition engine" is a technology that analyzes voice data and converts it into text data.

[0246] An "facial expression analysis engine" is a technology that analyzes a person's facial expressions from video data and identifies their emotional state.

[0247] "Speech content" is text information obtained by analyzing voice data.

[0248] A "speaker" is a person who speaks each utterance in the voice data.

[0249] "Speech duration" is data indicating the length of time that a particular speaker is speaking.

[0250] The "proportion of speaking time" is the proportion of the time a particular speaker is speaking to the total conversation time.

[0251] A "textualized conversation log" is data that has been recorded by converting voice data into text using a voice recognition engine.

[0252] An "emotion icon" is an icon that visually represents an emotional state identified by the facial expression analysis engine.

[0253] The "evaluation score" is a score obtained by evaluating the quality of speech and responses based on conversation logs and sentiment analysis data.

[0254] "Customer service" refers to the act of a store clerk providing products or services to customers.

[0255] "Customer emotion" refers to the emotional state that a customer displays while being served.

[0256] "Suggestion" is a function that suggests what to talk about next or what to do next.

[0257] The present invention relates to a system that supports customer service in brick-and-mortar stores by acquiring and analyzing voice and video data in real time. In particular, the use of smart glasses enables store clerks to communicate effectively with customers. This system is composed of the following elements: a means for acquiring voice and video data, a voice recognition engine, a facial expression analysis engine, a means for displaying emotional icons, a means for automatically converting conversation logs into text, a means for calculating speech time and evaluation scores, a suggestion function, and a means for accessing a dictionary database.

[0258] Setup and Connection

[0259] The server is connected to the smart glasses and provides tools to assist store clerks in serving customers. When a user (store clerk) starts serving customers, they put on the smart glasses and the system starts automatically, which starts data collection.

[0260] Acquiring audio and video data

[0261] The smart glasses, which are the terminals, use a microphone and camera to capture audio and video data in real time. This data is sent to the server in streaming format. The user simply needs to provide normal customer service without any awareness of the device.

[0262] Speech recognition and talking time percentage calculation

[0263] The server passes the voice data to a speech recognition engine to identify the content and speaker. The speech recognition engine analyzes the data to generate text data and record the content and speaker information. The server then aggregates the speaking time for each speaker and calculates the "talking time ratio" for the user and customer. This information is sent to the device in real time and overlaid on the smart glasses display.

[0264] sentiment analysis

[0265] The server passes the video data to a facial expression analysis engine to identify emotions from facial expressions, and also passes the audio data to a tone of voice analysis engine to identify emotions from tone of voice. These emotion data are integrated and displayed as emotion icons on the user's smart glasses.

[0266] Automatic text conversion of conversation logs

[0267] The server records the text data obtained from the speech recognition engine in real time as a conversation log, which is periodically sent to the terminal and displayed in real time on the user's smart glasses.

[0268] Calculation of evaluation score

[0269] The server evaluates the quality of the conversation and response based on the conversation log and sentiment analysis data, and calculates an evaluation score. This score is displayed in real time on the user's smart glasses, and suggests areas for improvement as needed.

[0270] Suggestion function

[0271] The server analyzes the conversation and progress between the user and the customer, and generates suggestions for next steps and solutions, which are displayed in real time as a pop-up on the user's smart glasses.

[0272] Quick Dictionary Function

[0273] When a user selects a word they don't understand, the device sends the word to the server, which searches for the word's meaning in a dictionary database and sends the results to the user's smart glasses, allowing the user to see the meaning of the selected word in a pop-up window.

[0274] Retrospective training function

[0275] The server stores all data (audio, video, conversation logs, sentiment analysis) during customer service and provides a review interface that users can access later. Based on the stored data, users can replay and analyze past customer service interactions and repeatedly practice with simulation videos that include areas for improvement.

[0276] Specific examples

[0277] For example, if a user serves a customer for 10 minutes, the device sends the audio and video data of the service to the server in real time. The server analyzes the audio data and obtains the results: the user spoke for 6 minutes and the customer spoke for 4 minutes. Based on this, the smart glasses display a speaking time ratio of 60% vs. 40%. At the same time, the content of the conversation is converted into text and displayed in real time on the user's smart glasses as a conversation log. The server also analyzes the customer's facial expressions and tone of voice, and displays emotion icons such as "interested" or "anxious."

[0278] Prompt Sentence Examples

[0279] Here are some example prompts to generate suggestions for the AI:

[0280] Conversation text:

[0281] Salesperson: Welcome. What item are you looking for today?

[0282] Customer: Yes, I'm looking for a gift for Ochugen.

[0283] Suggestion: Do you have a desired price range or a specific genre? You can also take a look at our catalog.

[0284] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0285] Step 1:

[0286] The smart glasses terminal captures audio and video data in real time using a microphone and camera from the moment the customer begins serving customers, and this data is sent to a server in streaming format.

[0287] Input: Real-time audio and video data during customer service

[0288] Output: Streamed audio and video data

[0289] Step 2:

[0290] The server passes the acquired voice data to a voice recognition engine, which analyzes and records the voice data as text data and speaker information. The voice recognition engine uses, for example, Google Speech Recognition.

[0291] Input: Streamed audio data

[0292] Data processing: Converting voice data into text (voice recognition)

[0293] Output: Text data and speaker information

[0294] Step 3:

[0295] At the same time, the server passes the video data to an expression analysis engine (such as DeepFace) to identify emotions from facial expressions, and a voice tone analysis engine to determine emotions from voice.

[0296] Input: Streamed video and audio data

[0297] Data processing: Facial expression analysis from video data, tone of voice analysis from audio data

[0298] Output: Emotion data (facial expressions and tone of voice)

[0299] Step 4:

[0300] The server records the text data obtained from the voice recognition engine as a conversation log, and transmits the textual conversation log to the terminal in real time at specific intervals.

[0301] Input: Text data

[0302] Data processing: Record text data as a conversation log

[0303] Output: Real-time conversation log

[0304] Step 5:

[0305] The server aggregates the speaking time of each speaker and calculates the percentage of time the user (store clerk) is talking and the percentage of time the customer is talking. This information is sent to the device in real time and displayed on the smart glasses display.

[0306] Input: Text data and speaker information

[0307] Data calculation: counting speech time and calculating percentage

[0308] Output: Percentage of speaking time

[0309] Step 6:

[0310] The server combines facial expression data and tone of voice data and sends them to the terminal as an emotional icon, allowing the user to visually grasp the customer's emotions.

[0311] Input: Emotion data (facial expressions and tone of voice)

[0312] Data processing: Emotion data integration and emotion icon generation

[0313] Output:Emotion icon

[0314] Step 7:

[0315] The server calculates an evaluation score based on the conversation log and emotion data and sends it to the device. The evaluation score is calculated based on factors such as the customer's reaction and the balance of their speech.

[0316] Input: Conversation logs and emotion data

[0317] Data calculation: Calculation of evaluation score

[0318] Output: Evaluation score

[0319] Step 8:

[0320] The server analyzes the conversation content and progress and generates suggestions for what to say next. The suggestions are displayed in real time in a pop-up format on the device.

[0321] Input: Conversation log and progress

[0322] Data Computation: Generating Suggestions with Generative AI Models

[0323] Output: Suggestion

[0324] Step 9:

[0325] When a user selects a word they do not understand, the device sends the word to the server, which searches for the meaning of the word in a dictionary database and sends the results to the device, where they are displayed in a pop-up window.

[0326] Input: selected word

[0327] Data processing: Retrieving meanings from dictionary databases

[0328] Output: Word meaning

[0329] Step 10:

[0330] The server stores all data (audio, video, conversation logs, and sentiment analysis data) during customer service and provides a review interface that users can access later. This allows users to replay and analyze past customer service interactions and repeatedly practice areas for improvement using simulation videos.

[0331] Input: Audio data, video data, conversation logs, sentiment analysis data

[0332] Data processing: storing data and providing an interface for reviewing it

[0333] Output: Retrospective interface

[0334] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0335] This invention relates to a multi-functional system that captures and analyzes audio and video data in real time to improve the quality of online business negotiations and interviews. This system provides more efficient communication support by combining it with an emotion engine that recognizes the user's emotions.

[0336] Explanation of program processing

[0337] Setup and Connection

[0338] The server connects the user's device to the main Web meeting product. When the user starts a Web meeting, the device notifies the server of the start of the session. Upon receiving this notification, the server sends the initial setup data for the tool to the user's device, completing the setup of the entire system.

[0339] Acquiring audio and video data

[0340] The device captures audio and video data in real time from its microphone and camera, and sends the captured data in streaming format to the server. The user simply conducts a regular web meeting.

[0341] Speech recognition and talking time percentage calculation

[0342] The server passes the received voice data to a voice recognition engine, which analyzes it to identify the content of the speech and the speaker. The analyzed text data and speaker information are recorded, and the server tallies the speaking time for each speaker and calculates the "percentage of speaking time." This information is sent to the device and displayed as an overlay on the user's screen.

[0343] sentiment analysis

[0344] The server passes video data to a facial expression analysis engine to identify emotions from facial expressions. It also passes audio data to a tone of voice analysis engine to identify emotions from vocal tones. In addition, the present invention incorporates an emotion engine that recognizes the user's emotions and integrates these data to generate an emotion icon. The generated emotion icon is displayed on the user's screen.

[0345] Automatic text conversion of conversation logs

[0346] The server records the conversation content obtained from the speech recognition engine as a text log in real time. The recorded text log is periodically sent to the terminal and displayed on the user's screen in real time.

[0347] Calculation of evaluation score

[0348] The server analyzes the conversation logs and sentiment analysis data to evaluate the user's performance based on best practices. The evaluation score is calculated in real time and sent to the user's device. Based on the evaluation score, the device presents improvements and suggestions to the user.

[0349] Quick Dictionary Function

[0350] When a user selects a word they do not understand, the device sends the word to the server, which searches for the meaning of the word in a dictionary database and sends the results back to the device, allowing the user to see the meaning of the selected word in a pop-up window.

[0351] Suggestion function

[0352] The server analyzes the conversation and its progress, and generates suggestions for what to say next and what questions to ask. These suggestions are sent to the user's device in real time and displayed in a pop-up format.

[0353] Retrospective training function

[0354] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and repeatedly practice areas where they failed or needed improvement using simulation videos.

[0355] Specific examples

[0356] When a user conducts a 20-minute online interview, the device captures audio and video data in real time and sends it to the server. The server analyzes the audio data and calculates the ratio of the user speaking for 12 minutes to the other person speaking for 8 minutes. This information is displayed on the screen. In addition, emotion icons derived from a facial expression analysis engine and a tone of voice analysis engine are displayed on the screen to help the user check for missed questions and consistency in the topic.

[0357] The present invention provides a multi-functional system that further improves the quality of communication by combining it with an emotion engine that recognizes the user's emotions, allowing users to efficiently conduct online business negotiations and online interviews.

[0358] The processing flow will be explained below.

[0359] Explanation of program processing

[0360] Setup and Connection

[0361] Step 1:

[0362] The server connects users' devices and major web meeting products.

[0363] Step 2:

[0364] When a user starts a Web meeting, the terminal notifies the server of the start of the session.

[0365] Step 3:

[0366] The server transmits the initial setting data of the tool to the user's terminal.

[0367] Acquiring audio and video data

[0368] Step 1:

[0369] The device captures audio and video data in real time from the microphone and camera.

[0370] Step 2:

[0371] The terminal transmits the acquired audio and video data to the server in streaming format.

[0372] Speech recognition and talking time percentage calculation

[0373] Step 1:

[0374] The server passes the received voice data to the voice recognition engine.

[0375] Step 2:

[0376] The voice recognition engine analyzes the voice data and identifies what is being said and who is speaking.

[0377] Step 3:

[0378] The server records the analyzed text data and speaker information.

[0379] Step 4:

[0380] The server aggregates the speaking time for each speaker and calculates the "proportion of speaking time" for the user and the other party.

[0381] Step 5:

[0382] The device overlays the percentage of speaking time received from the server on the screen.

[0383] sentiment analysis

[0384] Step 1:

[0385] The server passes the video data to the facial expression analysis engine.

[0386] Step 2:

[0387] The facial expression analysis engine analyzes video data and identifies emotions from facial expressions.

[0388] Step 3:

[0389] The server passes the audio data to a tone of voice analysis engine.

[0390] Step 4:

[0391] A voice tone analysis engine analyzes audio data and identifies emotions from the tone of voice.

[0392] Step 5:

[0393] The server also analyzes the user's facial expressions and voice and passes them to an emotion engine that recognizes the user's emotions.

[0394] Step 6:

[0395] The emotion engine identifies the user's emotion and integrates it with other emotion data.

[0396] Step 7:

[0397] The server generates an emotional icon from the integrated emotional data.

[0398] Step 8:

[0399] The device overlays the emotion icons received from the server on the screen.

[0400] Automatic text conversion of conversation logs

[0401] Step 1:

[0402] The server records the conversation content obtained from the voice recognition engine as a text log in real time.

[0403] Step 2:

[0404] The server periodically sends the text log to the user's terminal.

[0405] Step 3:

[0406] The device overlays the received conversation log on the screen in real time.

[0407] Calculation of evaluation score

[0408] Step 1:

[0409] The server analyzes the conversation logs and sentiment analysis data and evaluates them based on best practice rules.

[0410] Step 2:

[0411] The server calculates the evaluation score and sends the score to the user's terminal.

[0412] Step 3:

[0413] The device displays the received evaluation score on the screen and suggests areas for improvement if necessary.

[0414] Quick Dictionary Function

[0415] Step 1:

[0416] When the user selects a word they do not understand, the terminal sends the word to the server.

[0417] Step 2:

[0418] The server searches the dictionary database for the meaning of the word and sends the result to the terminal.

[0419] Step 3:

[0420] The device will pop up the meaning of the word on the screen.

[0421] Suggestion function

[0422] Step 1:

[0423] The server analyzes the conversation content and progress and generates suggestions for what to say next.

[0424] Step 2:

[0425] The server sends the generated suggestions to the user's terminal.

[0426] Step 3:

[0427] The device displays the received suggestions in a pop-up format on the screen.

[0428] Retrospective training function

[0429] Step 1:

[0430] The server stores all data from the meeting (audio, video, conversation logs, sentiment analysis).

[0431] Step 2:

[0432] The server provides a review interface that users can access after the meeting has ended.

[0433] Step 3:

[0434] Users can access the saved data to replay and analyze past business meetings and interviews.

[0435] Step 4:

[0436] Users can repeatedly practice the areas where they failed or the points they want to improve using simulation videos.

[0437] Specific examples

[0438] For example, if a user conducts a 20-minute online interview, the device captures audio and video data from the microphone and camera in real time and sends it in streaming format to the server. The server analyzes the audio data and calculates the percentage of time the user spoke for 12 minutes and the other party spoke for 8 minutes. This information is overlaid on the user's screen in the form of 68% vs. 32%.

[0439] The server then analyzes the user's facial expressions and tone of voice using an emotion engine, and combines the acquired emotion data to generate an emotion icon. For example, if the server determines that the user is nervous, it will display the emotion icon as "nervous."

[0440] The conversation log is converted into text in real time by a speech recognition engine and displayed on the user's screen. An evaluation score is also calculated in real time, allowing the user to instantly evaluate their own performance.

[0441] When a user uses the quick dictionary feature for a word they don't understand, the meaning of the selected word is displayed in a pop-up, and a suggestion function is also activated to keep the conversation flowing smoothly.

[0442] Finally, after the interview, you can use the retrospective training function to play back past interview data and repeatedly practice areas where you failed or need improvement using simulation footage.

[0443] Example 2

[0444] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0445] In conventional online business meetings and interviews, there was no system that could perform real-time emotion recognition, speech content analysis, evaluation score calculation, next topic suggestion, and even retrospective training all in one place. As a result, users had to rely on multiple tools and manual analysis, making it difficult to improve the quality of efficient communication. In addition, there was little real-time feedback, leading to missed opportunities for improvement.

[0446] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0447] In this invention, the server includes means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine and identifying the content and speaker of the speech, means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions, means for aggregating the speaking time for each speaker and calculating the percentage of speaking time, means for displaying a text-converted conversation log in real time, means for displaying the emotion analysis results as an emotion icon, means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data, means for analyzing the progress of the conversation and generating and displaying suggestions for what to say next, means for searching for and displaying the meaning of a selected word from a dictionary database, and means for unified management of voice, video, conversation log, and emotion analysis data and for replaying and analyzing past business negotiations and interviews based on the stored data. This allows users to receive real-time and comprehensive data analysis and feedback, significantly improving the quality of online business negotiations and online interviews.

[0448] "Audio data" refers to a speaker's audio signal collected through a microphone.

[0449] "Video data" refers to frames of video and their sequences collected through a camera.

[0450] A "voice recognition engine" refers to software or hardware for converting voice data into text data.

[0451] An "facial expression analysis engine" refers to software or hardware that analyzes human facial expressions from video data and identifies their emotions.

[0452] "Utterance content" refers to the specific words uttered by the speaker and their meaning.

[0453] "Speaker" refers to the person speaking.

[0454] "Speech time" refers to the total amount of time a speaker is actually speaking.

[0455] "Proportion of speaking time" refers to the ratio of the time each speaker is speaking to the total time.

[0456] A "conversation log" refers to a record of voice data converted into text.

[0457] "Emotion icons" refers to a display format that shows emotions obtained through facial expression analysis and voice analysis using shapes and emoticons.

[0458] "Evaluation score" refers to the result of quantifying a user's performance based on conversation logs and sentiment analysis data.

[0459] "Suggestions" refer to the next thing to say or questions that are generated based on the progress of the conversation.

[0460] A "dictionary database" refers to a database that associates words with their meanings.

[0461] "Stored Data" refers to audio, video, conversation logs, and sentiment analysis data collected during meetings.

[0462] A "natural language processing engine" refers to software or hardware for analyzing and generating text data.

[0463] This invention relates to a multi-functional system that captures and analyzes audio and video data in real time to improve the quality of online business negotiations and interviews. This system provides more efficient communication support by combining it with an emotion engine that recognizes the user's emotions.

[0464] The main components of this system include the user's device, server, voice recognition engine, facial expression analysis engine, voice tone analysis engine, emotion engine, dictionary database, and natural language processing engine. These elements work together to provide comprehensive meeting support functions.

[0465] The user's device captures audio and video data in real time using a built-in or externally connected microphone and camera. The captured data is sent to the server in a streaming format such as WebRTC. The user simply conducts a regular web meeting, and the data is analyzed by the server.

[0466] The server passes the audio data received in streaming format to a speech recognition engine such as Google Cloud Speech-to-Text for analysis. This identifies the content of the speech and the speaker, and the analyzed text data and speaker information are recorded in a database. The server also tallies the speaking time for each speaker and calculates the "percentage of speaking time." This information is sent to the user's device and displayed on the screen.

[0467] The server then passes the video data to Microsoft Azure's Cognitive Services Facial Expression Recognition API to identify emotions from facial expressions. It also uses IBM Watson's voice tone analysis to identify emotions from voice tone. The emotion engine combines these data to generate emotion icons, which are displayed on the user's screen in real time.

[0468] Furthermore, based on the conversation log and sentiment analysis data, the server evaluates the user's performance based on best practices. This evaluation score is also calculated in real time and sent to the user's device. Based on the evaluation score, the device presents the user with suggestions and areas for improvement.

[0469] The system also includes a quick dictionary function, which means that when a user selects a word they do not understand, the system sends the word to the server, searches for the meaning of the word in the dictionary database, and displays the results in a pop-up.

[0470] The server analyzes the conversation content and progress using a natural language processing (NLP) engine, and generates suggestions for what to say next and what questions to ask. These suggestions are also sent to the user's device in real time and displayed in a pop-up window.

[0471] Even after the meeting, the server stores the audio, video, conversation log, and sentiment analysis data, and provides a review interface that allows users to replay and analyze past business negotiations and interviews based on this data.Users can replay and analyze past business negotiations and interviews based on the saved data, and repeatedly practice areas where they failed or need improvement using simulation videos.

[0472] For example, if a user conducts a 20-minute online interview, the device captures audio and video data in real time and sends it to the server. The server analyzes the audio data and calculates the ratio of the user speaking for 12 minutes to the other person speaking for 8 minutes. This information is displayed on the screen. In addition, emotion icons obtained by a facial expression analysis engine and a tone of voice analysis engine are displayed on the screen to help the user check for missed questions and consistency in the topic.

[0473] An example of a prompt might be, "During the online interview, please tell us how the other person's emotions changed while they were listening to you."

[0474] As described above, the present invention is a system that provides real-time and comprehensive data analysis and feedback in online business negotiations and interviews, enabling users to efficiently improve the quality of their communication.

[0475] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0476] Step 1:

[0477] The device captures audio and video data in real time using a built-in or externally connected microphone and camera. The inputs are audio signals from the microphone and video data from the camera. These data are converted into a streaming format (such as H.264 or AAC) and sent to the server. Specifically, the device driver captures the data and transfers it to the server using the WebRTC protocol.

[0478] Step 2:

[0479] The server passes the voice data received from the device to a voice recognition engine such as Google Cloud Speech-to-Text. The input is streaming voice data. This is sent as an API request, and text data and speaker information are obtained as analysis results. Specifically, the server processes the voice data in batches and sends an API request to the voice recognition engine for each batch. The analyzed text data and speaker information are then recorded in a database.

[0480] Step 3:

[0481] The server passes the received video data to the Microsoft Azure Cognitive Services Facial Expression Recognition API to identify emotions. The input is real-time video data sent from the device. The facial expression analysis engine analyzes the video frames and outputs emotional data. Specifically, it sends an API request for each video frame, maps the obtained emotional data, and records it in a database.

[0482] Step 4:

[0483] The server passes the voice data to IBM Watson's tone of voice analysis engine, which identifies emotions from the tone of voice. The input is streaming voice data. The analysis results in emotional data based on the tone of voice. Specifically, the system sends the voice data in the form of an API request, receives the analyzed emotional data, and records it in a database.

[0484] Step 5:

[0485] The server integrates the emotion data obtained from the video and audio data to generate an emotion icon. The inputs are the results of facial expression analysis and tone of voice analysis. The emotion engine integrates these and outputs an emotion icon. Specifically, the integration algorithm is executed to generate an emotion icon. This icon is then sent to the user's device in real time.

[0486] Step 6:

[0487] The server records the text data obtained from the output of the speech recognition engine as a conversation log and sends it to the device in real time. The input is the analyzed text data. The log recorded in the database is sent to the device via WebSocket at regular intervals and displayed on the user's screen.

[0488] Step 7:

[0489] The server runs an algorithm that evaluates user performance based on conversation logs and sentiment analysis data. The inputs are textual conversation logs and emotional icon data. The server calculates an evaluation score and sends it to the device in real time. Specifically, the server uses a machine learning model to analyze the data and calculate a score.

[0490] Step 8:

[0491] When a user selects a word that the device does not understand, it sends the word to the server and searches for the meaning of the word in the dictionary database. The input is the word selected by the user. The meaning of the word retrieved from the server is displayed in a pop-up. Specifically, it captures the selection operation, sends an API request to the server, retrieves information from the dictionary database, and displays it.

[0492] Step 9:

[0493] The server analyzes the conversation content and progress, and generates the next thing to say and questions. The input is the real-time conversation log and progress data. The natural language processing (NLP) engine performs the analysis, generates suggestions, and sends them to the device. Specifically, the NLP engine analyzes the text data, creates a list of suggestions, and sends them in real time.

[0494] Step 10:

[0495] The server stores all data from the meeting (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. The input is all data collected during the meeting. This is stored in an integrated database so that users can access it later. Specifically, the server stores the data in cloud storage and provides an interface that users can use to play and analyze it through a browser.

[0496] (Application example 2)

[0497] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0498] Today, there are a wide variety of communication methods, but improving the quality of communication is especially important in online business meetings and interviews. However, conventional systems simply collect voice and video data and lack the support to identify emotions and promote effective communication. Furthermore, in situations such as customer support, effective dialogue between support staff and customers is important, and there is a growing need for systems that can analyze and support the progress of conversations in real time.

[0499] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0500] In this invention, the server includes: means for acquiring voice data and video data; means for passing the acquired voice data to a voice recognition engine to identify the content of the utterance and the speaker; means for passing the acquired video data to a facial expression analysis engine to identify emotions from facial expressions; means for aggregating the speaking time for each speaker and calculating the percentage of speaking time; means for displaying a text-converted conversation log in real time; means for displaying the emotion analysis results as emotion icons; means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data; means for analyzing the progress of the conversation and generating and displaying suggestions for what to say next; means for accessing stored data to play back and analyze past contact sessions or interviews; means for searching and displaying the meaning of selected words from a dictionary database; and means for a generative AI model to generate a prompt sentence to say next in real time using the emotion analysis results and the conversation log. This makes it possible to dramatically improve the quality of communication by identifying a user's emotions through real-time analysis of voice and video data and automatically suggesting what to say next.

[0501] "Voice data" is digital information that records the user's speech and surrounding sounds, and is acquired via a voice input device such as a microphone.

[0502] "Video data" is digital information that records visual information, including the movements of objects and people, acquired through a video input device such as a camera.

[0503] A "voice recognition engine" is a technology that analyzes input voice data and converts it into text data.

[0504] An "expression analysis engine" is a technology that analyzes the facial behavior and expressions of people in video data to determine their emotions.

[0505] The term "speaker" is a concept that refers to a specific person speaking in the audio data.

[0506] "Speaking time" refers to the total time that a particular speaker is speaking in the audio data.

[0507] A "textualized conversation log" is a record of the contents of voice data converted into text format.

[0508] An "emotion icon" is an icon that visually represents an emotion identified by the emotion analysis engine.

[0509] The "evaluation score" is a numerical evaluation of a user's performance based on conversation logs and sentiment analysis data.

[0510] "Conversation progress" is the process of analyzing the content of a conversation and its progress in real time.

[0511] "Suggestion" is a function that suggests the next topic or question to talk about based on the progress of the conversation.

[0512] "Stored Data" refers to all audio, video, conversation logs, and sentiment analysis data recorded by the system, including past contact sessions and interviews.

[0513] A "dictionary database" is a database that stores the meanings and related information of words.

[0514] A "generative AI model" is an algorithm generated using machine learning technology that generates the next prompt sentence to be spoken based on the conversation log and sentiment analysis results.

[0515] This invention relates to a multi-functional system that improves the quality of communication in business negotiations and interviews by identifying a user's emotions through real-time analysis of audio data and video data and automatically suggesting what to say next. Specific embodiments of this system will be described below.

[0516] Setup and Connection

[0517] The server connects the user's device to the main web meeting service. When a user starts a web meeting, the device notifies the server of the start of the session. Upon receiving this notification, the server sends the initial setup data for the tool to the user's device, completing the setup of the entire system.

[0518] Acquiring audio and video data

[0519] The device captures audio and video data in real time from its microphone and camera, and sends the captured data in streaming format to the server. The user simply conducts a regular web meeting.

[0520] Speech recognition and speaker identification

[0521] The server passes the received voice data to a voice recognition engine, which analyzes it to identify the content of the speech and the speaker. The analyzed text data and speaker information are recorded, and the server tallies the speaking time for each speaker and calculates the "percentage of speaking time." This information is sent to the device and displayed as an overlay on the user's screen.

[0522] sentiment analysis

[0523] The server passes video data to a facial expression analysis engine to identify emotions from facial expressions. It also passes audio data to a voice tone analysis engine to identify emotions from voice tone. Additionally, the server incorporates an emotion engine that recognizes the user's emotions and combines these data to generate an emotion icon. The generated emotion icon is then displayed on the user's screen.

[0524] Automatic text conversion of conversation logs

[0525] The server records the conversation content obtained from the speech recognition engine as a text log in real time. The recorded text log is periodically sent to the terminal and displayed on the user's screen in real time.

[0526] Calculation of evaluation score

[0527] The server analyzes the conversation logs and sentiment analysis data to evaluate the user's performance based on best practices. The evaluation score is calculated in real time and sent to the user's device. Based on the evaluation score, the device presents improvements and suggestions to the user.

[0528] Quick Dictionary Function

[0529] When a user selects a word they do not understand, the device sends the word to the server, which searches for the meaning of the word in a dictionary database and sends the results back to the device, allowing the user to see the meaning of the selected word in a pop-up window.

[0530] Suggestion function

[0531] The server analyzes the conversation content and progress, and generates suggestions for what to say next or what questions to ask. These suggestions are sent to the user's device in real time and displayed in a pop-up format. At this time, a generative AI model can be used to generate prompts. For example, a prompt such as "What should be the next question after this topic?" can be generated.

[0532] Retrospective training function

[0533] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and repeatedly practice areas where they failed or needed improvement using simulation videos.

[0534] Specific examples

[0535] A customer support agent conducts a 20-minute web chat session, during which the system analyzes audio and video data to identify agent and customer sentiment. Based on the conversation's progress, the system displays appropriate next questions and suggestions in real time, and provides data for retrospective training after the session ends. Using a generative AI model, the system can generate prompts such as, "What's the next question we should ask about this topic?"

[0536] As described above, this system can dramatically improve the quality of communication by identifying the user's emotions through real-time analysis of voice and video data and automatically suggesting what to say next.

[0537] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0538] Step 1: Setup and Connection

[0539] The server connects the user's terminal to the main Web meeting service. When a user starts a Web meeting, the terminal notifies the server of the start of the session. The input is the session start notification data, and the output is the initial setup data for the tool. The server receives this notification and sends the initial setup data for the tool to the user's terminal, completing the setup of the entire system.

[0540] Step 2: Acquiring audio and video data

[0541] The device captures audio and video data in real time from the microphone and camera. The captured data is sent to the server in streaming format as input data. The output is streaming audio and video data. The user simply needs to proceed with a normal web meeting.

[0542] Step 3: Speech recognition and speaker identification

[0543] The server passes the received voice data to a voice recognition engine for analysis. This analysis identifies the content of the speech and the speaker. The input is stream-format voice data, and the output is text data and speaker information. The server records this data and tallys up the speaking time for each speaker.

[0544] Step 4: Sentiment analysis

[0545] The server passes video data to a facial expression analysis engine to identify emotions from facial expressions. It also passes audio data to a tone of voice analysis engine to identify emotions from vocal tone. The input is streamed video and audio data, and the output is the emotion analysis results. These data are integrated to generate an emotional icon. The generated emotional icon is sent to the terminal and displayed on the user's screen.

[0546] Step 5: Automatic transcription of conversation logs

[0547] The server records the conversation content obtained from the speech recognition engine as a text log in real time. The input is text data and the output is a text log. This text log is periodically sent to the terminal and displayed on the user's screen in real time.

[0548] Step 6: Calculating the evaluation score

[0549] The server analyzes the conversation log and sentiment analysis data and evaluates the user's performance based on best practices. The input is the conversation log and sentiment analysis data, and the output is an evaluation score. The evaluation score is calculated in real time and sent to the user's device. The device then presents the user with suggestions and areas for improvement based on the evaluation score.

[0550] Step 7: Quick Dictionary Function

[0551] When the user selects a word that the terminal does not understand, it sends the word to the server. The input is the selected word, and the output is the meaning of the word. The server searches for the meaning of the word in the dictionary database and sends the result to the terminal. The user is prompted to confirm the meaning of the selected word in a pop-up window.

[0552] Step 8: Suggestions

[0553] The server analyzes the conversation content and progress, and generates suggestions for what to say next and what questions to ask. At this time, a generative AI model can be used to generate prompts. The input is the conversation log and progress, and the output is suggestions and prompts. These suggestions are sent to the device in real time and displayed in a pop-up format.

[0554] Step 9: Retrospective Training Function

[0555] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. The input is the meeting data, and the output is the review interface and past data. Based on the stored data, users can replay and analyze past business negotiations and interviews, and repeatedly practice using simulation videos to identify areas where they failed or need improvement.

[0556] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0557] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0558] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0559] [Second embodiment]

[0560] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0561] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0562] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0563] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0564] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0565] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0566] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0567] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0568] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0569] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0570] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0571] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0572] This invention relates to a system that supports online business negotiations and online interviews by acquiring and analyzing voice and video data in real time. This system is mainly composed of the following elements: a means for acquiring voice and video data, a voice recognition engine, a facial expression analysis engine, a means for displaying emotional icons, a means for automatically converting conversation logs into text, a means for calculating speech time and evaluation scores, a suggestion function, and a means for accessing a dictionary database.

[0573] Explanation of program processing

[0574] Setup and Connection

[0575] The server connects the user's device to a major web meeting product (such as a general web conferencing system). When the user starts a web meeting, the device connects to the server and launches a support tool, which starts collecting data for business negotiations or interviews.

[0576] Acquiring audio and video data

[0577] The device captures audio and video data in real time from the microphone and camera. This data is then sent directly to the server in streaming format. The user is then able to proceed with a normal web meeting without being aware of anything.

[0578] Speech recognition and talking time percentage calculation

[0579] The server passes the voice data to a voice recognition engine, which identifies the content of the speech and the speaker. The voice recognition engine analyzes the data, generates text data, and records the content of the speech and speaker information. The server then aggregates the speaking time for each speaker and calculates the "talking time ratio" for the user and the other party. This information is sent to the device in real time and displayed as an overlay on the screen.

[0580] sentiment analysis

[0581] The server passes the video data to a facial expression analysis engine, which identifies emotions from facial expressions. The audio data is also passed to a tone of voice analysis engine, which identifies emotions from the tone of voice. These emotion data are integrated and displayed on the user's screen as an emotion icon.

[0582] Automatic text conversion of conversation logs

[0583] The server records the text data obtained from the speech recognition engine in real time as a conversation log, which is periodically sent to the terminal and displayed on the user's screen in real time.

[0584] Calculation of evaluation score

[0585] The server evaluates the conversation logs and sentiment analysis data based on best practice rules to calculate a rating score, which is displayed in real time on the user's device and suggests areas for improvement if necessary.

[0586] Quick Dictionary Function

[0587] When a user selects a word they don't understand, the device sends it to the server, which searches for the word's meaning in the dictionary database and sends the results to the user's device, where the user can see the meaning of the selected word in a pop-up window.

[0588] Suggestion function

[0589] The server analyzes the conversation and its progress, and generates suggestions for what to say next and what questions to ask. These suggestions are sent to the user's device in real time and displayed in a pop-up window.

[0590] Retrospective training function

[0591] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and practice repeatedly using simulation videos that include areas of failure and areas for improvement.

[0592] Specific examples

[0593] For example, consider a case where a user has a 10-minute online business meeting. The device sends the audio and video data of the meeting to the server in real time. The server analyzes the audio data and obtains the results: the user spoke for 6 minutes and the other party spoke for 4 minutes. Based on this, the speaking time ratios of 68% vs. 32% are displayed on the screen. At the same time, the content of the conversation is converted into text and displayed in real time on the user's screen as a conversation log. The server also analyzes the other party's facial expressions and tone of voice, and displays emotion icons such as "skepticism" or "interest."

[0594] In this way, the present invention is a multifunctional, real-time support system that enables users to conduct online business negotiations and online interviews more effectively.

[0595] The processing flow will be explained below.

[0596] Explanation of program processing

[0597] Setup and Connection

[0598] Step 1:

[0599] The server connects users' devices and major web meeting products.

[0600] Step 2:

[0601] When a user starts a Web meeting, the terminal notifies the server of the start of the session.

[0602] Step 3:

[0603] The server transmits the initial setting data of the tool to the user's terminal.

[0604] Acquiring audio and video data

[0605] Step 1:

[0606] The device captures audio and video data in real time from the microphone and camera.

[0607] Step 2:

[0608] The terminal transmits the acquired audio and video data to the server in streaming format.

[0609] Speech recognition and talking time percentage calculation

[0610] Step 1:

[0611] The server passes the received voice data to the voice recognition engine.

[0612] Step 2:

[0613] The voice recognition engine analyzes the voice data and identifies what is being said and who is speaking.

[0614] Step 3:

[0615] The server records the analyzed text data and speaker information.

[0616] Step 4:

[0617] The server aggregates the speaking time for each speaker and calculates the "proportion of speaking time" for the user and the other party.

[0618] Step 5:

[0619] The device overlays the percentage of speaking time received from the server on the screen.

[0620] sentiment analysis

[0621] Step 1:

[0622] The server passes the video data to the facial expression analysis engine.

[0623] Step 2:

[0624] The facial expression analysis engine analyzes video data and identifies emotions from facial expressions.

[0625] Step 3:

[0626] The server passes the audio data to a tone of voice analysis engine.

[0627] Step 4:

[0628] A voice tone analysis engine analyzes audio data and identifies emotions from the tone of voice.

[0629] Step 5:

[0630] The server integrates the obtained emotion data and generates an emotion icon.

[0631] Step 6:

[0632] The device overlays the emotion icons received from the server on the screen.

[0633] Automatic text conversion of conversation logs

[0634] Step 1:

[0635] The server records the conversation content obtained from the voice recognition engine as a text log in real time.

[0636] Step 2:

[0637] The server periodically sends the text log to the user's terminal.

[0638] Step 3:

[0639] The device overlays the received conversation log on the screen in real time.

[0640] Calculation of evaluation score

[0641] Step 1:

[0642] The server analyzes the conversation logs and sentiment analysis data and evaluates them based on best practice rules.

[0643] Step 2:

[0644] The server calculates the evaluation score and sends the score to the user's terminal.

[0645] Step 3:

[0646] The device displays the received evaluation score on the screen and suggests areas for improvement if necessary.

[0647] Quick Dictionary Function

[0648] Step 1:

[0649] When the user selects a word they do not understand, the terminal sends the word to the server.

[0650] Step 2:

[0651] The server searches the dictionary database for the meaning of the word and sends the result to the terminal.

[0652] Step 3:

[0653] The device will pop up the meaning of the word on the screen.

[0654] Suggestion function

[0655] Step 1:

[0656] The server analyzes the conversation content and progress and generates suggestions for what to say next.

[0657] Step 2:

[0658] The server sends the generated suggestions to the user's terminal.

[0659] Step 3:

[0660] The device displays the received suggestions in a pop-up format on the screen.

[0661] Retrospective training function

[0662] Step 1:

[0663] The server stores all data from the meeting (audio, video, conversation logs, sentiment analysis).

[0664] Step 2:

[0665] The server provides a review interface that users can access after the meeting has ended.

[0666] Step 3:

[0667] Users can access the saved data to replay and analyze past business meetings and interviews.

[0668] Step 4:

[0669] Users can repeatedly practice the areas where they failed or the points they want to improve using simulation videos.

[0670] Example 1

[0671] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0672] In recent years, remote work and online business negotiations and interviews have become more common, but communication in these online environments is considered more difficult than face-to-face communication. It is particularly difficult to grasp the other person's reactions and emotions in real time, making it difficult to effectively progress and evaluate the dialogue. Therefore, there is a demand for a system that supports online business negotiations and interviews, with a function that can analyze the other person's reactions and emotions in real time and enable users to communicate effectively.

[0673] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0674] In this invention, the server includes means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine and identifying the content of the speech and the speaker, means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions, means for aggregating the speaking time for each speaker and calculating the percentage of speaking time, means for displaying a textualized conversation log in real time, means for displaying the emotion analysis results as emotion icons, means for recording the text data acquired from the voice recognition engine as a conversation log and transmitting it to a terminal, and means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data. This makes it possible to grasp the other party's reactions and emotions in real time and provide an environment in which users can communicate effectively online.

[0675] "Audio data" refers to sound information acquired through a microphone during online business negotiations or online interviews.

[0676] "Video data" refers to video information captured through a camera during online business negotiations or interviews.

[0677] A "voice recognition engine" is a general term for technology that automatically identifies the content and speaker of an utterance from acquired voice data and converts it into text data.

[0678] "Facial expression analysis engine" is a general term for technology that analyzes facial expressions from acquired video data and identifies emotions from them.

[0679] "Speech duration" refers to the length of time during which a particular speaker speaks.

[0680] "Text conversion" is the process of analyzing audio data and recording what is being said as written information.

[0681] A "conversation log" is a record of text data generated by a voice recognition engine, recorded in chronological order.

[0682] "Emotion icons" are shapes or facial expression icons that visually represent the results of the emotion analysis engine.

[0683] The "evaluation score" is a numerical evaluation value that represents the quality and effectiveness of a conversation based on conversation logs and sentiment analysis data.

[0684] "Suggestion" is a function that analyzes the content and progress of a conversation and suggests what to talk about next or what questions to ask.

[0685] A "dictionary database" is a database that stores the meanings and explanations of words.

[0686] "Real-time display" refers to the act of displaying acquired data and analysis results instantly with almost no delay.

[0687] "Cloud storage" is an online storage service for storing data over the Internet.

[0688] This invention is a system that supports online business negotiations and online interviews by capturing and analyzing audio and video data in real time. This system has the following configuration, including the following main hardware and software elements: a microphone, a camera, a voice recognition engine, a facial expression analysis engine, a suggestion function, and a dictionary database access means.

[0689] Setup and Connection

[0690] The server connects the user's device to the main web meeting product. When the user starts a web meeting, the device connects to the server and launches the support tool, which starts collecting data for the business meeting or interview.

[0691] Acquiring audio and video data

[0692] The device uses a microphone and camera to capture audio and video data in real time, which is then sent in streaming format to a server for analysis.

[0693] Speech recognition and talking time percentage calculation

[0694] The server passes the voice data to a voice recognition engine, which identifies the content of the speech and the speaker. The voice recognition engine converts the voice data into text data and records the content of the speech and speaker information. The server then aggregates the speaking time for each speaker and calculates the "talking time ratio" for the user and the other party. This information is sent to the device in real time and displayed as an overlay on the screen.

[0695] As a specific example of use, the voice recognition engine uses the Google Cloud Speech-to-Text API to convert actual voice data into text with extremely high accuracy.

[0696] sentiment analysis

[0697] The server passes the video data to a facial expression analysis engine, which identifies emotions from facial expressions. The audio data is passed to a tone of voice analysis engine, which identifies emotions from the tone of voice. These emotion data are integrated and displayed on the user's screen as an emotion icon.

[0698] As a specific example of use, the facial expression analysis engine uses Microsoft Azure Face API, and the voice tone analysis engine uses IBM Watson Tone Analyzer.

[0699] Automatic text conversion of conversation logs

[0700] The server records the text data obtained from the speech recognition engine in real time as a conversation log, which is periodically sent to the terminal and displayed on the user's screen in real time.

[0701] Calculation of evaluation score

[0702] The server evaluates the conversation log and sentiment analysis data using best practice rules to calculate a score, which is displayed on the device in real time and suggests areas for improvement if necessary.

[0703] Quick Dictionary Function

[0704] When a user selects a word they do not understand, the device sends it to the server, which searches for the meaning of the word in a dictionary database and sends the results to the device, allowing the user to check the meaning of the selected word in a pop-up window.

[0705] Suggestion function

[0706] The server analyzes the conversation and progress, and generates suggestions for what to say next and what questions to ask. These suggestions are sent to the device in real time and displayed in a pop-up format.

[0707] Retrospective training function

[0708] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and practice repeatedly using simulation videos that include areas of failure and areas for improvement.

[0709] Specific examples

[0710] For example, if a user has a 10-minute online business meeting, the device sends the audio and video data of the meeting to the server in real time. The server analyzes the audio data and determines that the user spoke for 6 minutes and the other party spoke for 4 minutes. Based on this, the speaking time ratio of 68% vs. 32% is displayed on the screen. At the same time, the content of the conversation is converted into text and displayed on the user's screen in real time as a conversation log. The server also analyzes the other party's facial expressions and tone of voice, and displays emotion icons such as "suspicious" and "interested."

[0711] Prompt Sentence Examples

[0712] "Please explain about a system that allows users to check emotional icons and conversation logs in real time during online business negotiations and suggests ways to improve their conversation skills based on the other party's reactions."

[0713] In this way, the present invention is a multi-functional, real-time support system that enables users to conduct online business negotiations and online interviews more effectively.

[0714] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0715] Step 1: Setup and Connection

[0716] The server connects the user's device to the main Web meeting product. When the user starts a Web meeting, the device works with the server to launch the support tool.

[0717] Specific behavior:

[0718] Input: User action to start a web meeting

[0719] Data processing: The server connects to the meeting using the API of the web conferencing system.

[0720] Output: Gets meeting session information and sends it to the device

[0721] Step 2: Acquiring audio and video data

[0722] The device uses a microphone and camera to capture audio and video data in real time, and transmits this data to a server in streaming format.

[0723] Specific behavior:

[0724] Input: Audio data (from microphone), video data (from camera)

[0725] Data Processing: Real-time acquisition and synchronization of audio and video data

[0726] Output: Sends audio and video data to the server in streaming format

[0727] Step 3: Speech recognition and speaking time calculation

[0728] The server passes the voice data to a speech recognition engine to identify the content and speaker, then aggregates the speaking time for each speaker and calculates the "speaking time percentage."

[0729] Specific behavior:

[0730] Input: A stream of audio data

[0731] Data calculation: Converting voice data into text using the Google Cloud Speech-to-Text API, and identifying the content and speaker

[0732] Output: Send the aggregated results to the terminal and overlay them on the screen

[0733] Step 4: Sentiment analysis

[0734] The server passes the video data to an expression analysis engine to determine emotions from facial expressions, and passes the audio data to a tone of voice analysis engine to determine emotions.

[0735] Specific behavior:

[0736] Input: Video data (to facial expression analysis engine), audio data (to voice tone analysis engine)

[0737] Data calculation: Facial expression analysis using Microsoft Azure Face API, voice tone analysis using IBM Watson Tone Analyzer

[0738] Output: Generate emoticons and send them to the device

[0739] Step 5: Automatic transcription of conversation logs

[0740] The server records the text data obtained from the voice recognition engine as a conversation log and periodically sends it to the terminal.

[0741] Specific behavior:

[0742] Input: Text data from the speech recognition engine

[0743] Data processing: Recorded in a database as a conversation log

[0744] Output: Real-time updated conversation log sent to device

[0745] Step 6: Calculating the evaluation score

[0746] The server calculates a rating score using best practice rules based on conversation logs and sentiment analysis data.

[0747] Specific behavior:

[0748] Input: Conversation logs, sentiment analysis data

[0749] Data calculations: evaluated based on best practice rules

[0750] Output: Send the evaluation score to the terminal and display it.

[0751] Step 7: Quick Dictionary Function

[0752] The terminal sends the word selected by the user to the server, which searches for the meaning of the word from a dictionary database and returns it.

[0753] Specific behavior:

[0754] Input: Selected word (from terminal)

[0755] Data operations: Retrieving word meanings from dictionary databases (e.g., Merriam-Webster API)

[0756] Output: Send word meanings to terminal and display

[0757] Step 8: Suggestions

[0758] Based on the analysis of the conversation content and the progress, the server generates the next topic to talk about and questions, and sends them to the user's device in real time.

[0759] Specific behavior:

[0760] Input: Conversation content and progress data

[0761] Data Computation: Using generative AI models to generate suggestions

[0762] Output: Send suggestions to the device and display them in a popup

[0763] Step 9: Retrospective Training Function

[0764] The server stores all data from the meeting and allows users to access a review interface after the meeting, allowing them to replay and analyze past data and identify areas for improvement.

[0765] Specific behavior:

[0766] Input: Meeting audio, video, conversation logs, and sentiment analysis data

[0767] Data processing: Save data to cloud storage

[0768] Output: Provides a retrospective interface that allows users to play back and analyze data

[0769] (Application example 1)

[0770] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0771] Conventional customer service systems lack real-time feedback and support to enable store clerks to communicate smoothly with customers. This can lead to an inability to respond quickly to customer needs, which can lead to a decline in customer satisfaction. This can also have a negative impact on store sales and repeat customer rates. The present invention aims to solve these problems and provide a system that enables store clerks to communicate more effectively and smoothly with customers.

[0772] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0773] In this invention, the server includes means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine and identifying the content of the utterance and the speaker, means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions, means for aggregating the speaking time for each speaker and calculating the percentage of speaking time, means for displaying a text-converted conversation log in real time, means for displaying the emotion analysis results as an emotion icon, means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data, and means for analyzing customer emotions and conversation content during customer service in real time and generating and displaying suggestions for what to say next. This enables store clerks to analyze communication with customers in real time and appropriately determine the optimal response and next action, thereby improving customer satisfaction.

[0774] "Audio data" refers to audio information recorded in digital format.

[0775] "Video data" refers to video information recorded in digital format.

[0776] A "voice recognition engine" is a technology that analyzes voice data and converts it into text data.

[0777] An "facial expression analysis engine" is a technology that analyzes a person's facial expressions from video data and identifies their emotional state.

[0778] "Speech content" is text information obtained by analyzing voice data.

[0779] A "speaker" is a person who speaks each utterance in the voice data.

[0780] "Speech duration" is data indicating the length of time that a particular speaker is speaking.

[0781] The "proportion of speaking time" is the proportion of the time a particular speaker is speaking to the total conversation time.

[0782] A "textualized conversation log" is data that has been recorded by converting voice data into text using a voice recognition engine.

[0783] An "emotion icon" is an icon that visually represents an emotional state identified by the facial expression analysis engine.

[0784] The "evaluation score" is a score obtained by evaluating the quality of speech and responses based on conversation logs and sentiment analysis data.

[0785] "Customer service" refers to the act of a store clerk providing products or services to customers.

[0786] "Customer emotion" refers to the emotional state that a customer displays while being served.

[0787] "Suggestion" is a function that suggests what to talk about next or what to do next.

[0788] The present invention relates to a system that supports customer service in brick-and-mortar stores by acquiring and analyzing voice and video data in real time. In particular, the use of smart glasses enables store clerks to communicate effectively with customers. This system is composed of the following elements: a means for acquiring voice and video data, a voice recognition engine, a facial expression analysis engine, a means for displaying emotional icons, a means for automatically converting conversation logs into text, a means for calculating speech time and evaluation scores, a suggestion function, and a means for accessing a dictionary database.

[0789] Setup and Connection

[0790] The server is connected to the smart glasses and provides tools to assist store clerks in serving customers. When a user (store clerk) starts serving customers, they put on the smart glasses and the system starts automatically, which starts data collection.

[0791] Acquiring audio and video data

[0792] The smart glasses, which are the terminals, use a microphone and camera to capture audio and video data in real time. This data is sent to the server in streaming format. The user simply needs to provide normal customer service without any awareness of the device.

[0793] Speech recognition and talking time percentage calculation

[0794] The server passes the voice data to a speech recognition engine to identify the content and speaker. The speech recognition engine analyzes the data to generate text data and record the content and speaker information. The server then aggregates the speaking time for each speaker and calculates the "talking time ratio" for the user and customer. This information is sent to the device in real time and overlaid on the smart glasses display.

[0795] sentiment analysis

[0796] The server passes the video data to a facial expression analysis engine to identify emotions from facial expressions, and also passes the audio data to a tone of voice analysis engine to identify emotions from tone of voice. These emotion data are integrated and displayed as emotion icons on the user's smart glasses.

[0797] Automatic text conversion of conversation logs

[0798] The server records the text data obtained from the speech recognition engine in real time as a conversation log, which is periodically sent to the terminal and displayed in real time on the user's smart glasses.

[0799] Calculation of evaluation score

[0800] The server evaluates the quality of the conversation and response based on the conversation log and sentiment analysis data, and calculates an evaluation score. This score is displayed in real time on the user's smart glasses, and suggests areas for improvement as needed.

[0801] Suggestion function

[0802] The server analyzes the conversation and progress between the user and the customer, and generates suggestions for next steps and solutions, which are displayed in real time as a pop-up on the user's smart glasses.

[0803] Quick Dictionary Function

[0804] When a user selects a word they don't understand, the device sends the word to the server, which searches for the word's meaning in a dictionary database and sends the results to the user's smart glasses, allowing the user to see the meaning of the selected word in a pop-up window.

[0805] Retrospective training function

[0806] The server stores all data (audio, video, conversation logs, sentiment analysis) during customer service and provides a review interface that users can access later. Based on the stored data, users can replay and analyze past customer service interactions and repeatedly practice with simulation videos that include areas for improvement.

[0807] Specific examples

[0808] For example, if a user serves a customer for 10 minutes, the device sends the audio and video data of the service to the server in real time. The server analyzes the audio data and obtains the results: the user spoke for 6 minutes and the customer spoke for 4 minutes. Based on this, the smart glasses display a speaking time ratio of 60% vs. 40%. At the same time, the content of the conversation is converted into text and displayed in real time on the user's smart glasses as a conversation log. The server also analyzes the customer's facial expressions and tone of voice, and displays emotion icons such as "interested" or "anxious."

[0809] Prompt Sentence Examples

[0810] Here are some example prompts to generate suggestions for the AI:

[0811] Conversation text:

[0812] Salesperson: Welcome. What item are you looking for today?

[0813] Customer: Yes, I'm looking for a gift for Ochugen.

[0814] Suggestion: Do you have a desired price range or a specific genre? You can also take a look at our catalog.

[0815] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0816] Step 1:

[0817] The smart glasses terminal captures audio and video data in real time using a microphone and camera from the moment the customer begins serving customers, and this data is sent to a server in streaming format.

[0818] Input: Real-time audio and video data during customer service

[0819] Output: Streamed audio and video data

[0820] Step 2:

[0821] The server passes the acquired voice data to a voice recognition engine, which analyzes and records the voice data as text data and speaker information. The voice recognition engine uses, for example, Google Speech Recognition.

[0822] Input: Streamed audio data

[0823] Data processing: Converting voice data into text (voice recognition)

[0824] Output: Text data and speaker information

[0825] Step 3:

[0826] At the same time, the server passes the video data to an expression analysis engine (such as DeepFace) to identify emotions from facial expressions, and a voice tone analysis engine to determine emotions from voice.

[0827] Input: Streamed video and audio data

[0828] Data processing: Facial expression analysis from video data, tone of voice analysis from audio data

[0829] Output: Emotion data (facial expressions and tone of voice)

[0830] Step 4:

[0831] The server records the text data obtained from the voice recognition engine as a conversation log, and transmits the textual conversation log to the terminal in real time at specific intervals.

[0832] Input: Text data

[0833] Data processing: Record text data as a conversation log

[0834] Output: Real-time conversation log

[0835] Step 5:

[0836] The server aggregates the speaking time of each speaker and calculates the percentage of time the user (store clerk) is talking and the percentage of time the customer is talking. This information is sent to the device in real time and displayed on the smart glasses display.

[0837] Input: Text data and speaker information

[0838] Data calculation: counting speech time and calculating percentage

[0839] Output: Percentage of speaking time

[0840] Step 6:

[0841] The server combines facial expression data and tone of voice data and sends them to the terminal as an emotional icon, allowing the user to visually grasp the customer's emotions.

[0842] Input: Emotion data (facial expressions and tone of voice)

[0843] Data processing: Emotion data integration and emotion icon generation

[0844] Output:Emotion icon

[0845] Step 7:

[0846] The server calculates an evaluation score based on the conversation log and emotion data and sends it to the device. The evaluation score is calculated based on factors such as the customer's reaction and the balance of their speech.

[0847] Input: Conversation logs and emotion data

[0848] Data calculation: Calculation of evaluation score

[0849] Output: Evaluation score

[0850] Step 8:

[0851] The server analyzes the conversation content and progress and generates suggestions for what to say next. The suggestions are displayed in real time in a pop-up format on the device.

[0852] Input: Conversation log and progress

[0853] Data Computation: Generating Suggestions with Generative AI Models

[0854] Output: Suggestion

[0855] Step 9:

[0856] When a user selects a word they do not understand, the device sends the word to the server, which searches for the meaning of the word in a dictionary database and sends the results to the device, where they are displayed in a pop-up window.

[0857] Input: selected word

[0858] Data processing: Retrieving meanings from dictionary databases

[0859] Output: Word meaning

[0860] Step 10:

[0861] The server stores all data (audio, video, conversation logs, and sentiment analysis data) during customer service and provides a review interface that users can access later. This allows users to replay and analyze past customer service interactions and repeatedly practice areas for improvement using simulation videos.

[0862] Input: Audio data, video data, conversation logs, sentiment analysis data

[0863] Data processing: storing data and providing an interface for reviewing it

[0864] Output: Retrospective interface

[0865] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0866] This invention relates to a multi-functional system that captures and analyzes audio and video data in real time to improve the quality of online business negotiations and interviews. This system provides more efficient communication support by combining it with an emotion engine that recognizes the user's emotions.

[0867] Explanation of program processing

[0868] Setup and Connection

[0869] The server connects the user's device to the main Web meeting product. When the user starts a Web meeting, the device notifies the server of the start of the session. Upon receiving this notification, the server sends the initial setup data for the tool to the user's device, completing the setup of the entire system.

[0870] Acquiring audio and video data

[0871] The device captures audio and video data in real time from its microphone and camera, and sends the captured data in streaming format to the server. The user simply conducts a regular web meeting.

[0872] Speech recognition and talking time percentage calculation

[0873] The server passes the received voice data to a voice recognition engine, which analyzes it to identify the content of the speech and the speaker. The analyzed text data and speaker information are recorded, and the server tallies the speaking time for each speaker and calculates the "percentage of speaking time." This information is sent to the device and displayed as an overlay on the user's screen.

[0874] sentiment analysis

[0875] The server passes video data to a facial expression analysis engine to identify emotions from facial expressions. It also passes audio data to a tone of voice analysis engine to identify emotions from vocal tones. In addition, the present invention incorporates an emotion engine that recognizes the user's emotions and integrates these data to generate an emotion icon. The generated emotion icon is displayed on the user's screen.

[0876] Automatic text conversion of conversation logs

[0877] The server records the conversation content obtained from the speech recognition engine as a text log in real time. The recorded text log is periodically sent to the terminal and displayed on the user's screen in real time.

[0878] Calculation of evaluation score

[0879] The server analyzes the conversation logs and sentiment analysis data to evaluate the user's performance based on best practices. The evaluation score is calculated in real time and sent to the user's device. Based on the evaluation score, the device presents improvements and suggestions to the user.

[0880] Quick Dictionary Function

[0881] When a user selects a word they do not understand, the device sends the word to the server, which searches for the meaning of the word in a dictionary database and sends the results back to the device, allowing the user to see the meaning of the selected word in a pop-up window.

[0882] Suggestion function

[0883] The server analyzes the conversation and its progress, and generates suggestions for what to say next and what questions to ask. These suggestions are sent to the user's device in real time and displayed in a pop-up format.

[0884] Retrospective training function

[0885] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and repeatedly practice areas where they failed or needed improvement using simulation videos.

[0886] Specific examples

[0887] When a user conducts a 20-minute online interview, the device captures audio and video data in real time and sends it to the server. The server analyzes the audio data and calculates the ratio of the user speaking for 12 minutes to the other person speaking for 8 minutes. This information is displayed on the screen. In addition, emotion icons derived from a facial expression analysis engine and a tone of voice analysis engine are displayed on the screen to help the user check for missed questions and consistency in the topic.

[0888] The present invention provides a multi-functional system that further improves the quality of communication by combining it with an emotion engine that recognizes the user's emotions, allowing users to efficiently conduct online business negotiations and online interviews.

[0889] The processing flow will be explained below.

[0890] Explanation of program processing

[0891] Setup and Connection

[0892] Step 1:

[0893] The server connects users' devices and major web meeting products.

[0894] Step 2:

[0895] When a user starts a Web meeting, the terminal notifies the server of the start of the session.

[0896] Step 3:

[0897] The server transmits the initial setting data of the tool to the user's terminal.

[0898] Acquiring audio and video data

[0899] Step 1:

[0900] The device captures audio and video data in real time from the microphone and camera.

[0901] Step 2:

[0902] The terminal transmits the acquired audio and video data to the server in streaming format.

[0903] Speech recognition and talking time percentage calculation

[0904] Step 1:

[0905] The server passes the received voice data to the voice recognition engine.

[0906] Step 2:

[0907] The voice recognition engine analyzes the voice data and identifies what is being said and who is speaking.

[0908] Step 3:

[0909] The server records the analyzed text data and speaker information.

[0910] Step 4:

[0911] The server aggregates the speaking time for each speaker and calculates the "proportion of speaking time" for the user and the other party.

[0912] Step 5:

[0913] The device overlays the percentage of speaking time received from the server on the screen.

[0914] sentiment analysis

[0915] Step 1:

[0916] The server passes the video data to the facial expression analysis engine.

[0917] Step 2:

[0918] The facial expression analysis engine analyzes video data and identifies emotions from facial expressions.

[0919] Step 3:

[0920] The server passes the audio data to a tone of voice analysis engine.

[0921] Step 4:

[0922] A voice tone analysis engine analyzes audio data and identifies emotions from the tone of voice.

[0923] Step 5:

[0924] The server also analyzes the user's facial expressions and voice and passes them to an emotion engine that recognizes the user's emotions.

[0925] Step 6:

[0926] The emotion engine identifies the user's emotion and integrates it with other emotion data.

[0927] Step 7:

[0928] The server generates an emotional icon from the integrated emotional data.

[0929] Step 8:

[0930] The device overlays the emotion icons received from the server on the screen.

[0931] Automatic text conversion of conversation logs

[0932] Step 1:

[0933] The server records the conversation content obtained from the voice recognition engine as a text log in real time.

[0934] Step 2:

[0935] The server periodically sends the text log to the user's terminal.

[0936] Step 3:

[0937] The device overlays the received conversation log on the screen in real time.

[0938] Calculation of evaluation score

[0939] Step 1:

[0940] The server analyzes the conversation logs and sentiment analysis data and evaluates them based on best practice rules.

[0941] Step 2:

[0942] The server calculates the evaluation score and sends the score to the user's terminal.

[0943] Step 3:

[0944] The device displays the received evaluation score on the screen and suggests areas for improvement if necessary.

[0945] Quick Dictionary Function

[0946] Step 1:

[0947] When the user selects a word they do not understand, the terminal sends the word to the server.

[0948] Step 2:

[0949] The server searches the dictionary database for the meaning of the word and sends the result to the terminal.

[0950] Step 3:

[0951] The device will pop up the meaning of the word on the screen.

[0952] Suggestion function

[0953] Step 1:

[0954] The server analyzes the conversation content and progress and generates suggestions for what to say next.

[0955] Step 2:

[0956] The server sends the generated suggestions to the user's terminal.

[0957] Step 3:

[0958] The device displays the received suggestions in a pop-up format on the screen.

[0959] Retrospective training function

[0960] Step 1:

[0961] The server stores all data from the meeting (audio, video, conversation logs, sentiment analysis).

[0962] Step 2:

[0963] The server provides a review interface that users can access after the meeting has ended.

[0964] Step 3:

[0965] Users can access the saved data to replay and analyze past business meetings and interviews.

[0966] Step 4:

[0967] Users can repeatedly practice the areas where they failed or the points they want to improve using simulation videos.

[0968] Specific examples

[0969] For example, if a user conducts a 20-minute online interview, the device captures audio and video data from the microphone and camera in real time and sends it in streaming format to the server. The server analyzes the audio data and calculates the percentage of time the user spoke for 12 minutes and the other party spoke for 8 minutes. This information is overlaid on the user's screen in the form of 68% vs. 32%.

[0970] The server then analyzes the user's facial expressions and tone of voice using an emotion engine, and combines the acquired emotion data to generate an emotion icon. For example, if the server determines that the user is nervous, it will display the emotion icon as "nervous."

[0971] The conversation log is converted into text in real time by a speech recognition engine and displayed on the user's screen. An evaluation score is also calculated in real time, allowing the user to instantly evaluate their own performance.

[0972] When a user uses the quick dictionary feature for a word they don't understand, the meaning of the selected word is displayed in a pop-up, and a suggestion function is also activated to keep the conversation flowing smoothly.

[0973] Finally, after the interview, you can use the retrospective training function to play back past interview data and repeatedly practice areas where you failed or need improvement using simulation footage.

[0974] Example 2

[0975] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0976] In conventional online business meetings and interviews, there was no system that could perform real-time emotion recognition, speech content analysis, evaluation score calculation, next topic suggestion, and even retrospective training all in one place. As a result, users had to rely on multiple tools and manual analysis, making it difficult to improve the quality of efficient communication. In addition, there was little real-time feedback, leading to missed opportunities for improvement.

[0977] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0978] In this invention, the server includes means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine and identifying the content and speaker of the speech, means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions, means for aggregating the speaking time for each speaker and calculating the percentage of speaking time, means for displaying a text-converted conversation log in real time, means for displaying the emotion analysis results as an emotion icon, means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data, means for analyzing the progress of the conversation and generating and displaying suggestions for what to say next, means for searching for and displaying the meaning of a selected word from a dictionary database, and means for unified management of voice, video, conversation log, and emotion analysis data and for replaying and analyzing past business negotiations and interviews based on the stored data. This allows users to receive real-time and comprehensive data analysis and feedback, significantly improving the quality of online business negotiations and online interviews.

[0979] "Audio data" refers to a speaker's audio signal collected through a microphone.

[0980] "Video data" refers to frames of video and their sequences collected through a camera.

[0981] A "voice recognition engine" refers to software or hardware for converting voice data into text data.

[0982] An "facial expression analysis engine" refers to software or hardware that analyzes human facial expressions from video data and identifies their emotions.

[0983] "Utterance content" refers to the specific words uttered by the speaker and their meaning.

[0984] "Speaker" refers to the person speaking.

[0985] "Speech time" refers to the total amount of time a speaker is actually speaking.

[0986] "Proportion of speaking time" refers to the ratio of the time each speaker is speaking to the total time.

[0987] A "conversation log" refers to a record of voice data converted into text.

[0988] "Emotion icons" refers to a display format that shows emotions obtained through facial expression analysis and voice analysis using shapes and emoticons.

[0989] "Evaluation score" refers to the result of quantifying a user's performance based on conversation logs and sentiment analysis data.

[0990] "Suggestions" refer to the next thing to say or questions that are generated based on the progress of the conversation.

[0991] A "dictionary database" refers to a database that associates words with their meanings.

[0992] "Stored Data" refers to audio, video, conversation logs, and sentiment analysis data collected during meetings.

[0993] A "natural language processing engine" refers to software or hardware for analyzing and generating text data.

[0994] This invention relates to a multi-functional system that captures and analyzes audio and video data in real time to improve the quality of online business negotiations and interviews. This system provides more efficient communication support by combining it with an emotion engine that recognizes the user's emotions.

[0995] The main components of this system include the user's device, server, voice recognition engine, facial expression analysis engine, voice tone analysis engine, emotion engine, dictionary database, and natural language processing engine. These elements work together to provide comprehensive meeting support functions.

[0996] The user's device captures audio and video data in real time using a built-in or externally connected microphone and camera. The captured data is sent to the server in a streaming format such as WebRTC. The user simply conducts a regular web meeting, and the data is analyzed by the server.

[0997] The server passes the audio data received in streaming format to a speech recognition engine such as Google Cloud Speech-to-Text for analysis. This identifies the content of the speech and the speaker, and the analyzed text data and speaker information are recorded in a database. The server also tallies the speaking time for each speaker and calculates the "percentage of speaking time." This information is sent to the user's device and displayed on the screen.

[0998] The server then passes the video data to Microsoft Azure's Cognitive Services Facial Expression Recognition API to identify emotions from facial expressions. It also uses IBM Watson's voice tone analysis to identify emotions from voice tone. The emotion engine combines these data to generate emotion icons, which are displayed on the user's screen in real time.

[0999] Furthermore, based on the conversation log and sentiment analysis data, the server evaluates the user's performance based on best practices. This evaluation score is also calculated in real time and sent to the user's device. Based on the evaluation score, the device presents the user with suggestions and areas for improvement.

[1000] The system also includes a quick dictionary function, which means that when a user selects a word they do not understand, the system sends the word to the server, searches for the meaning of the word in the dictionary database, and displays the results in a pop-up.

[1001] The server analyzes the conversation content and progress using a natural language processing (NLP) engine, and generates suggestions for what to say next and what questions to ask. These suggestions are also sent to the user's device in real time and displayed in a pop-up window.

[1002] Even after the meeting, the server stores the audio, video, conversation log, and sentiment analysis data, and provides a review interface that allows users to replay and analyze past business negotiations and interviews based on this data.Users can replay and analyze past business negotiations and interviews based on the saved data, and repeatedly practice areas where they failed or need improvement using simulation videos.

[1003] For example, if a user conducts a 20-minute online interview, the device captures audio and video data in real time and sends it to the server. The server analyzes the audio data and calculates the ratio of the user speaking for 12 minutes to the other person speaking for 8 minutes. This information is displayed on the screen. In addition, emotion icons obtained by a facial expression analysis engine and a tone of voice analysis engine are displayed on the screen to help the user check for missed questions and consistency in the topic.

[1004] An example of a prompt might be, "During the online interview, please tell us how the other person's emotions changed while they were listening to you."

[1005] As described above, the present invention is a system that provides real-time and comprehensive data analysis and feedback in online business negotiations and interviews, enabling users to efficiently improve the quality of their communication.

[1006] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1007] Step 1:

[1008] The device captures audio and video data in real time using a built-in or externally connected microphone and camera. The inputs are audio signals from the microphone and video data from the camera. These data are converted into a streaming format (such as H.264 or AAC) and sent to the server. Specifically, the device driver captures the data and transfers it to the server using the WebRTC protocol.

[1009] Step 2:

[1010] The server passes the voice data received from the device to a voice recognition engine such as Google Cloud Speech-to-Text. The input is streaming voice data. This is sent as an API request, and text data and speaker information are obtained as analysis results. Specifically, the server processes the voice data in batches and sends an API request to the voice recognition engine for each batch. The analyzed text data and speaker information are then recorded in a database.

[1011] Step 3:

[1012] The server passes the received video data to the Microsoft Azure Cognitive Services Facial Expression Recognition API to identify emotions. The input is real-time video data sent from the device. The facial expression analysis engine analyzes the video frames and outputs emotional data. Specifically, it sends an API request for each video frame, maps the obtained emotional data, and records it in a database.

[1013] Step 4:

[1014] The server passes the voice data to IBM Watson's tone of voice analysis engine, which identifies emotions from the tone of voice. The input is streaming voice data. The analysis results in emotional data based on the tone of voice. Specifically, the system sends the voice data in the form of an API request, receives the analyzed emotional data, and records it in a database.

[1015] Step 5:

[1016] The server integrates the emotion data obtained from the video and audio data to generate an emotion icon. The inputs are the results of facial expression analysis and tone of voice analysis. The emotion engine integrates these and outputs an emotion icon. Specifically, the integration algorithm is executed to generate an emotion icon. This icon is then sent to the user's device in real time.

[1017] Step 6:

[1018] The server records the text data obtained from the output of the speech recognition engine as a conversation log and sends it to the device in real time. The input is the analyzed text data. The log recorded in the database is sent to the device via WebSocket at regular intervals and displayed on the user's screen.

[1019] Step 7:

[1020] The server runs an algorithm that evaluates user performance based on conversation logs and sentiment analysis data. The inputs are textual conversation logs and emotional icon data. The server calculates an evaluation score and sends it to the device in real time. Specifically, the server uses a machine learning model to analyze the data and calculate a score.

[1021] Step 8:

[1022] When a user selects a word that the device does not understand, it sends the word to the server and searches for the meaning of the word in the dictionary database. The input is the word selected by the user. The meaning of the word retrieved from the server is displayed in a pop-up. Specifically, it captures the selection operation, sends an API request to the server, retrieves information from the dictionary database, and displays it.

[1023] Step 9:

[1024] The server analyzes the conversation content and progress, and generates the next thing to say and questions. The input is the real-time conversation log and progress data. The natural language processing (NLP) engine performs the analysis, generates suggestions, and sends them to the device. Specifically, the NLP engine analyzes the text data, creates a list of suggestions, and sends them in real time.

[1025] Step 10:

[1026] The server stores all data from the meeting (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. The input is all data collected during the meeting. This is stored in an integrated database so that users can access it later. Specifically, the server stores the data in cloud storage and provides an interface that users can use to play and analyze it through a browser.

[1027] (Application example 2)

[1028] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1029] Today, there are a wide variety of communication methods, but improving the quality of communication is especially important in online business meetings and interviews. However, conventional systems simply collect voice and video data and lack the support to identify emotions and promote effective communication. Furthermore, in situations such as customer support, effective dialogue between support staff and customers is important, and there is a growing need for systems that can analyze and support the progress of conversations in real time.

[1030] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1031] In this invention, the server includes: means for acquiring voice data and video data; means for passing the acquired voice data to a voice recognition engine to identify the content of the utterance and the speaker; means for passing the acquired video data to a facial expression analysis engine to identify emotions from facial expressions; means for aggregating the speaking time for each speaker and calculating the percentage of speaking time; means for displaying a text-converted conversation log in real time; means for displaying the emotion analysis results as emotion icons; means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data; means for analyzing the progress of the conversation and generating and displaying suggestions for what to say next; means for accessing stored data to play back and analyze past contact sessions or interviews; means for searching and displaying the meaning of selected words from a dictionary database; and means for a generative AI model to generate a prompt sentence to say next in real time using the emotion analysis results and the conversation log. This makes it possible to dramatically improve the quality of communication by identifying a user's emotions through real-time analysis of voice and video data and automatically suggesting what to say next.

[1032] "Voice data" is digital information that records the user's speech and surrounding sounds, and is acquired via a voice input device such as a microphone.

[1033] "Video data" is digital information that records visual information, including the movements of objects and people, acquired through a video input device such as a camera.

[1034] A "voice recognition engine" is a technology that analyzes input voice data and converts it into text data.

[1035] An "expression analysis engine" is a technology that analyzes the facial behavior and expressions of people in video data to determine their emotions.

[1036] The term "speaker" is a concept that refers to a specific person speaking in the audio data.

[1037] "Speaking time" refers to the total time that a particular speaker is speaking in the audio data.

[1038] A "textualized conversation log" is a record of the contents of voice data converted into text format.

[1039] An "emotion icon" is an icon that visually represents an emotion identified by the emotion analysis engine.

[1040] The "evaluation score" is a numerical evaluation of a user's performance based on conversation logs and sentiment analysis data.

[1041] "Conversation progress" is the process of analyzing the content of a conversation and its progress in real time.

[1042] "Suggestion" is a function that suggests the next topic or question to talk about based on the progress of the conversation.

[1043] "Stored Data" refers to all audio, video, conversation logs, and sentiment analysis data recorded by the system, including past contact sessions and interviews.

[1044] A "dictionary database" is a database that stores the meanings and related information of words.

[1045] A "generative AI model" is an algorithm generated using machine learning technology that generates the next prompt sentence to be spoken based on the conversation log and sentiment analysis results.

[1046] This invention relates to a multi-functional system that improves the quality of communication in business negotiations and interviews by identifying a user's emotions through real-time analysis of audio data and video data and automatically suggesting what to say next. Specific embodiments of this system will be described below.

[1047] Setup and Connection

[1048] The server connects the user's device to the main web meeting service. When a user starts a web meeting, the device notifies the server of the start of the session. Upon receiving this notification, the server sends the initial setup data for the tool to the user's device, completing the setup of the entire system.

[1049] Acquiring audio and video data

[1050] The device captures audio and video data in real time from its microphone and camera, and sends the captured data in streaming format to the server. The user simply conducts a regular web meeting.

[1051] Speech recognition and speaker identification

[1052] The server passes the received voice data to a voice recognition engine, which analyzes it to identify the content of the speech and the speaker. The analyzed text data and speaker information are recorded, and the server tallies the speaking time for each speaker and calculates the "percentage of speaking time." This information is sent to the device and displayed as an overlay on the user's screen.

[1053] sentiment analysis

[1054] The server passes video data to a facial expression analysis engine to identify emotions from facial expressions. It also passes audio data to a voice tone analysis engine to identify emotions from voice tone. Additionally, the server incorporates an emotion engine that recognizes the user's emotions and combines these data to generate an emotion icon. The generated emotion icon is then displayed on the user's screen.

[1055] Automatic text conversion of conversation logs

[1056] The server records the conversation content obtained from the speech recognition engine as a text log in real time. The recorded text log is periodically sent to the terminal and displayed on the user's screen in real time.

[1057] Calculation of evaluation score

[1058] The server analyzes the conversation logs and sentiment analysis data to evaluate the user's performance based on best practices. The evaluation score is calculated in real time and sent to the user's device. Based on the evaluation score, the device presents improvements and suggestions to the user.

[1059] Quick Dictionary Function

[1060] When a user selects a word they do not understand, the device sends the word to the server, which searches for the meaning of the word in a dictionary database and sends the results back to the device, allowing the user to see the meaning of the selected word in a pop-up window.

[1061] Suggestion function

[1062] The server analyzes the conversation content and progress, and generates suggestions for what to say next or what questions to ask. These suggestions are sent to the user's device in real time and displayed in a pop-up format. At this time, a generative AI model can be used to generate prompts. For example, a prompt such as "What should be the next question after this topic?" can be generated.

[1063] Retrospective training function

[1064] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and repeatedly practice areas where they failed or needed improvement using simulation videos.

[1065] Specific examples

[1066] A customer support agent conducts a 20-minute web chat session, during which the system analyzes audio and video data to identify agent and customer sentiment. Based on the conversation's progress, the system displays appropriate next questions and suggestions in real time, and provides data for retrospective training after the session ends. Using a generative AI model, the system can generate prompts such as, "What's the next question we should ask about this topic?"

[1067] As described above, this system can dramatically improve the quality of communication by identifying the user's emotions through real-time analysis of voice and video data and automatically suggesting what to say next.

[1068] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1069] Step 1: Setup and Connection

[1070] The server connects the user's terminal to the main Web meeting service. When a user starts a Web meeting, the terminal notifies the server of the start of the session. The input is the session start notification data, and the output is the initial setup data for the tool. The server receives this notification and sends the initial setup data for the tool to the user's terminal, completing the setup of the entire system.

[1071] Step 2: Acquiring audio and video data

[1072] The device captures audio and video data in real time from the microphone and camera. The captured data is sent to the server in streaming format as input data. The output is streaming audio and video data. The user simply needs to proceed with a normal web meeting.

[1073] Step 3: Speech recognition and speaker identification

[1074] The server passes the received voice data to a voice recognition engine for analysis. This analysis identifies the content of the speech and the speaker. The input is stream-format voice data, and the output is text data and speaker information. The server records this data and tallys up the speaking time for each speaker.

[1075] Step 4: Sentiment analysis

[1076] The server passes video data to a facial expression analysis engine to identify emotions from facial expressions. It also passes audio data to a tone of voice analysis engine to identify emotions from vocal tone. The input is streamed video and audio data, and the output is the emotion analysis results. These data are integrated to generate an emotional icon. The generated emotional icon is sent to the terminal and displayed on the user's screen.

[1077] Step 5: Automatic transcription of conversation logs

[1078] The server records the conversation content obtained from the speech recognition engine as a text log in real time. The input is text data and the output is a text log. This text log is periodically sent to the terminal and displayed on the user's screen in real time.

[1079] Step 6: Calculating the evaluation score

[1080] The server analyzes the conversation log and sentiment analysis data and evaluates the user's performance based on best practices. The input is the conversation log and sentiment analysis data, and the output is an evaluation score. The evaluation score is calculated in real time and sent to the user's device. The device then presents the user with suggestions and areas for improvement based on the evaluation score.

[1081] Step 7: Quick Dictionary Function

[1082] When the user selects a word that the terminal does not understand, it sends the word to the server. The input is the selected word, and the output is the meaning of the word. The server searches for the meaning of the word in the dictionary database and sends the result to the terminal. The user is prompted to confirm the meaning of the selected word in a pop-up window.

[1083] Step 8: Suggestions

[1084] The server analyzes the conversation content and progress, and generates suggestions for what to say next and what questions to ask. At this time, a generative AI model can be used to generate prompts. The input is the conversation log and progress, and the output is suggestions and prompts. These suggestions are sent to the device in real time and displayed in a pop-up format.

[1085] Step 9: Retrospective Training Function

[1086] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. The input is the meeting data, and the output is the review interface and past data. Based on the stored data, users can replay and analyze past business negotiations and interviews, and repeatedly practice using simulation videos to identify areas where they failed or need improvement.

[1087] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1088] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1089] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1090] [Third embodiment]

[1091] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1092] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1093] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1094] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1095] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1096] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1097] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1098] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1099] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1100] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1101] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1102] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1103] This invention relates to a system that supports online business negotiations and online interviews by acquiring and analyzing voice and video data in real time. This system is mainly composed of the following elements: a means for acquiring voice and video data, a voice recognition engine, a facial expression analysis engine, a means for displaying emotional icons, a means for automatically converting conversation logs into text, a means for calculating speech time and evaluation scores, a suggestion function, and a means for accessing a dictionary database.

[1104] Explanation of program processing

[1105] Setup and Connection

[1106] The server connects the user's device to a major web meeting product (such as a general web conferencing system). When the user starts a web meeting, the device connects to the server and launches a support tool, which starts collecting data for business negotiations or interviews.

[1107] Acquiring audio and video data

[1108] The device captures audio and video data in real time from the microphone and camera. This data is then sent directly to the server in streaming format. The user is then able to proceed with a normal web meeting without being aware of anything.

[1109] Speech recognition and talking time percentage calculation

[1110] The server passes the voice data to a voice recognition engine, which identifies the content of the speech and the speaker. The voice recognition engine analyzes the data, generates text data, and records the content of the speech and speaker information. The server then aggregates the speaking time for each speaker and calculates the "talking time ratio" for the user and the other party. This information is sent to the device in real time and displayed as an overlay on the screen.

[1111] sentiment analysis

[1112] The server passes the video data to a facial expression analysis engine, which identifies emotions from facial expressions. The audio data is also passed to a tone of voice analysis engine, which identifies emotions from the tone of voice. These emotion data are integrated and displayed on the user's screen as an emotion icon.

[1113] Automatic text conversion of conversation logs

[1114] The server records the text data obtained from the speech recognition engine in real time as a conversation log, which is periodically sent to the terminal and displayed on the user's screen in real time.

[1115] Calculation of evaluation score

[1116] The server evaluates the conversation logs and sentiment analysis data based on best practice rules to calculate a rating score, which is displayed in real time on the user's device and suggests areas for improvement if necessary.

[1117] Quick Dictionary Function

[1118] When a user selects a word they don't understand, the device sends it to the server, which searches for the word's meaning in the dictionary database and sends the results to the user's device, where the user can see the meaning of the selected word in a pop-up window.

[1119] Suggestion function

[1120] The server analyzes the conversation and its progress, and generates suggestions for what to say next and what questions to ask. These suggestions are sent to the user's device in real time and displayed in a pop-up window.

[1121] Retrospective training function

[1122] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and practice repeatedly using simulation videos that include areas of failure and areas for improvement.

[1123] Specific examples

[1124] For example, consider a case where a user has a 10-minute online business meeting. The device sends the audio and video data of the meeting to the server in real time. The server analyzes the audio data and obtains the results: the user spoke for 6 minutes and the other party spoke for 4 minutes. Based on this, the speaking time ratios of 68% vs. 32% are displayed on the screen. At the same time, the content of the conversation is converted into text and displayed in real time on the user's screen as a conversation log. The server also analyzes the other party's facial expressions and tone of voice, and displays emotion icons such as "skepticism" or "interest."

[1125] In this way, the present invention is a multifunctional, real-time support system that enables users to conduct online business negotiations and online interviews more effectively.

[1126] The processing flow will be explained below.

[1127] Explanation of program processing

[1128] Setup and Connection

[1129] Step 1:

[1130] The server connects users' devices and major web meeting products.

[1131] Step 2:

[1132] When a user starts a Web meeting, the terminal notifies the server of the start of the session.

[1133] Step 3:

[1134] The server transmits the initial setting data of the tool to the user's terminal.

[1135] Acquiring audio and video data

[1136] Step 1:

[1137] The device captures audio and video data in real time from the microphone and camera.

[1138] Step 2:

[1139] The terminal transmits the acquired audio and video data to the server in streaming format.

[1140] Speech recognition and talking time percentage calculation

[1141] Step 1:

[1142] The server passes the received voice data to the voice recognition engine.

[1143] Step 2:

[1144] The voice recognition engine analyzes the voice data and identifies what is being said and who is speaking.

[1145] Step 3:

[1146] The server records the analyzed text data and speaker information.

[1147] Step 4:

[1148] The server aggregates the speaking time for each speaker and calculates the "proportion of speaking time" for the user and the other party.

[1149] Step 5:

[1150] The device overlays the percentage of speaking time received from the server on the screen.

[1151] sentiment analysis

[1152] Step 1:

[1153] The server passes the video data to the facial expression analysis engine.

[1154] Step 2:

[1155] The facial expression analysis engine analyzes video data and identifies emotions from facial expressions.

[1156] Step 3:

[1157] The server passes the audio data to a tone of voice analysis engine.

[1158] Step 4:

[1159] A voice tone analysis engine analyzes audio data and identifies emotions from the tone of voice.

[1160] Step 5:

[1161] The server integrates the obtained emotion data and generates an emotion icon.

[1162] Step 6:

[1163] The device overlays the emotion icons received from the server on the screen.

[1164] Automatic text conversion of conversation logs

[1165] Step 1:

[1166] The server records the conversation content obtained from the voice recognition engine as a text log in real time.

[1167] Step 2:

[1168] The server periodically sends the text log to the user's terminal.

[1169] Step 3:

[1170] The device overlays the received conversation log on the screen in real time.

[1171] Calculation of evaluation score

[1172] Step 1:

[1173] The server analyzes the conversation logs and sentiment analysis data and evaluates them based on best practice rules.

[1174] Step 2:

[1175] The server calculates the evaluation score and sends the score to the user's terminal.

[1176] Step 3:

[1177] The device displays the received evaluation score on the screen and suggests areas for improvement if necessary.

[1178] Quick Dictionary Function

[1179] Step 1:

[1180] When the user selects a word they do not understand, the terminal sends the word to the server.

[1181] Step 2:

[1182] The server searches the dictionary database for the meaning of the word and sends the result to the terminal.

[1183] Step 3:

[1184] The device will pop up the meaning of the word on the screen.

[1185] Suggestion function

[1186] Step 1:

[1187] The server analyzes the conversation content and progress and generates suggestions for what to say next.

[1188] Step 2:

[1189] The server sends the generated suggestions to the user's terminal.

[1190] Step 3:

[1191] The device displays the received suggestions in a pop-up format on the screen.

[1192] Retrospective training function

[1193] Step 1:

[1194] The server stores all data from the meeting (audio, video, conversation logs, sentiment analysis).

[1195] Step 2:

[1196] The server provides a review interface that users can access after the meeting has ended.

[1197] Step 3:

[1198] Users can access the saved data to replay and analyze past business meetings and interviews.

[1199] Step 4:

[1200] Users can repeatedly practice the areas where they failed or the points they want to improve using simulation videos.

[1201] Example 1

[1202] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1203] In recent years, remote work and online business negotiations and interviews have become more common, but communication in these online environments is considered more difficult than face-to-face communication. It is particularly difficult to grasp the other person's reactions and emotions in real time, making it difficult to effectively progress and evaluate the dialogue. Therefore, there is a demand for a system that supports online business negotiations and interviews, with a function that can analyze the other person's reactions and emotions in real time and enable users to communicate effectively.

[1204] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1205] In this invention, the server includes means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine and identifying the content of the speech and the speaker, means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions, means for aggregating the speaking time for each speaker and calculating the percentage of speaking time, means for displaying a textualized conversation log in real time, means for displaying the emotion analysis results as emotion icons, means for recording the text data acquired from the voice recognition engine as a conversation log and transmitting it to a terminal, and means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data. This makes it possible to grasp the other party's reactions and emotions in real time and provide an environment in which users can communicate effectively online.

[1206] "Audio data" refers to sound information acquired through a microphone during online business negotiations or online interviews.

[1207] "Video data" refers to video information captured through a camera during online business negotiations or interviews.

[1208] A "voice recognition engine" is a general term for technology that automatically identifies the content and speaker of an utterance from acquired voice data and converts it into text data.

[1209] "Facial expression analysis engine" is a general term for technology that analyzes facial expressions from acquired video data and identifies emotions from them.

[1210] "Speech duration" refers to the length of time during which a particular speaker speaks.

[1211] "Text conversion" is the process of analyzing audio data and recording what is being said as written information.

[1212] A "conversation log" is a record of text data generated by a voice recognition engine, recorded in chronological order.

[1213] "Emotion icons" are shapes or facial expression icons that visually represent the results of the emotion analysis engine.

[1214] The "evaluation score" is a numerical evaluation value that represents the quality and effectiveness of a conversation based on conversation logs and sentiment analysis data.

[1215] "Suggestion" is a function that analyzes the content and progress of a conversation and suggests what to talk about next or what questions to ask.

[1216] A "dictionary database" is a database that stores the meanings and explanations of words.

[1217] "Real-time display" refers to the act of displaying acquired data and analysis results instantly with almost no delay.

[1218] "Cloud storage" is an online storage service for storing data over the Internet.

[1219] This invention is a system that supports online business negotiations and online interviews by capturing and analyzing audio and video data in real time. This system has the following configuration, including the following main hardware and software elements: a microphone, a camera, a voice recognition engine, a facial expression analysis engine, a suggestion function, and a dictionary database access means.

[1220] Setup and Connection

[1221] The server connects the user's device to the main web meeting product. When the user starts a web meeting, the device connects to the server and launches the support tool, which starts collecting data for the business meeting or interview.

[1222] Acquiring audio and video data

[1223] The device uses a microphone and camera to capture audio and video data in real time, which is then sent in streaming format to a server for analysis.

[1224] Speech recognition and talking time percentage calculation

[1225] The server passes the voice data to a voice recognition engine, which identifies the content of the speech and the speaker. The voice recognition engine converts the voice data into text data and records the content of the speech and speaker information. The server then aggregates the speaking time for each speaker and calculates the "talking time ratio" for the user and the other party. This information is sent to the device in real time and displayed as an overlay on the screen.

[1226] As a specific example of use, the voice recognition engine uses the Google Cloud Speech-to-Text API to convert actual voice data into text with extremely high accuracy.

[1227] sentiment analysis

[1228] The server passes the video data to a facial expression analysis engine, which identifies emotions from facial expressions. The audio data is passed to a tone of voice analysis engine, which identifies emotions from the tone of voice. These emotion data are integrated and displayed on the user's screen as an emotion icon.

[1229] As a specific example of use, the facial expression analysis engine uses Microsoft Azure Face API, and the voice tone analysis engine uses IBM Watson Tone Analyzer.

[1230] Automatic text conversion of conversation logs

[1231] The server records the text data obtained from the speech recognition engine in real time as a conversation log, which is periodically sent to the terminal and displayed on the user's screen in real time.

[1232] Calculation of evaluation score

[1233] The server evaluates the conversation log and sentiment analysis data using best practice rules to calculate a score, which is displayed on the device in real time and suggests areas for improvement if necessary.

[1234] Quick Dictionary Function

[1235] When a user selects a word they do not understand, the device sends it to the server, which searches for the meaning of the word in a dictionary database and sends the results to the device, allowing the user to check the meaning of the selected word in a pop-up window.

[1236] Suggestion function

[1237] The server analyzes the conversation and progress, and generates suggestions for what to say next and what questions to ask. These suggestions are sent to the device in real time and displayed in a pop-up format.

[1238] Retrospective training function

[1239] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and practice repeatedly using simulation videos that include areas of failure and areas for improvement.

[1240] Specific examples

[1241] For example, if a user has a 10-minute online business meeting, the device sends the audio and video data of the meeting to the server in real time. The server analyzes the audio data and determines that the user spoke for 6 minutes and the other party spoke for 4 minutes. Based on this, the speaking time ratio of 68% vs. 32% is displayed on the screen. At the same time, the content of the conversation is converted into text and displayed on the user's screen in real time as a conversation log. The server also analyzes the other party's facial expressions and tone of voice, and displays emotion icons such as "suspicious" and "interested."

[1242] Prompt Sentence Examples

[1243] "Please explain about a system that allows users to check emotional icons and conversation logs in real time during online business negotiations and suggests ways to improve their conversation skills based on the other party's reactions."

[1244] In this way, the present invention is a multi-functional, real-time support system that enables users to conduct online business negotiations and online interviews more effectively.

[1245] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1246] Step 1: Setup and Connection

[1247] The server connects the user's device to the main Web meeting product. When the user starts a Web meeting, the device works with the server to launch the support tool.

[1248] Specific behavior:

[1249] Input: User action to start a web meeting

[1250] Data processing: The server connects to the meeting using the API of the web conferencing system.

[1251] Output: Gets meeting session information and sends it to the device

[1252] Step 2: Acquiring audio and video data

[1253] The device uses a microphone and camera to capture audio and video data in real time, and transmits this data to a server in streaming format.

[1254] Specific behavior:

[1255] Input: Audio data (from microphone), video data (from camera)

[1256] Data Processing: Real-time acquisition and synchronization of audio and video data

[1257] Output: Sends audio and video data to the server in streaming format

[1258] Step 3: Speech recognition and speaking time calculation

[1259] The server passes the voice data to a speech recognition engine to identify the content and speaker, then aggregates the speaking time for each speaker and calculates the "speaking time percentage."

[1260] Specific behavior:

[1261] Input: A stream of audio data

[1262] Data calculation: Converting voice data into text using the Google Cloud Speech-to-Text API, and identifying the content and speaker

[1263] Output: Send the aggregated results to the terminal and overlay them on the screen

[1264] Step 4: Sentiment analysis

[1265] The server passes the video data to an expression analysis engine to determine emotions from facial expressions, and passes the audio data to a tone of voice analysis engine to determine emotions.

[1266] Specific behavior:

[1267] Input: Video data (to facial expression analysis engine), audio data (to voice tone analysis engine)

[1268] Data calculation: Facial expression analysis using Microsoft Azure Face API, voice tone analysis using IBM Watson Tone Analyzer

[1269] Output: Generate emoticons and send them to the device

[1270] Step 5: Automatic transcription of conversation logs

[1271] The server records the text data obtained from the voice recognition engine as a conversation log and periodically sends it to the terminal.

[1272] Specific behavior:

[1273] Input: Text data from the speech recognition engine

[1274] Data processing: Recorded in a database as a conversation log

[1275] Output: Real-time updated conversation log sent to device

[1276] Step 6: Calculating the evaluation score

[1277] The server calculates a rating score using best practice rules based on conversation logs and sentiment analysis data.

[1278] Specific behavior:

[1279] Input: Conversation logs, sentiment analysis data

[1280] Data calculations: evaluated based on best practice rules

[1281] Output: Send the evaluation score to the terminal and display it.

[1282] Step 7: Quick Dictionary Function

[1283] The terminal sends the word selected by the user to the server, which searches for the meaning of the word from a dictionary database and returns it.

[1284] Specific behavior:

[1285] Input: Selected word (from terminal)

[1286] Data operations: Retrieving word meanings from dictionary databases (e.g., Merriam-Webster API)

[1287] Output: Send word meanings to terminal and display

[1288] Step 8: Suggestions

[1289] Based on the analysis of the conversation content and the progress, the server generates the next topic to talk about and questions, and sends them to the user's device in real time.

[1290] Specific behavior:

[1291] Input: Conversation content and progress data

[1292] Data Computation: Using generative AI models to generate suggestions

[1293] Output: Send suggestions to the device and display them in a popup

[1294] Step 9: Retrospective Training Function

[1295] The server stores all data from the meeting and allows users to access a review interface after the meeting, allowing them to replay and analyze past data and identify areas for improvement.

[1296] Specific behavior:

[1297] Input: Meeting audio, video, conversation logs, and sentiment analysis data

[1298] Data processing: Save data to cloud storage

[1299] Output: Provides a retrospective interface that allows users to play back and analyze data

[1300] (Application example 1)

[1301] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1302] Conventional customer service systems lack real-time feedback and support to enable store clerks to communicate smoothly with customers. This can lead to an inability to respond quickly to customer needs, which can lead to a decline in customer satisfaction. This can also have a negative impact on store sales and repeat customer rates. The present invention aims to solve these problems and provide a system that enables store clerks to communicate more effectively and smoothly with customers.

[1303] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1304] In this invention, the server includes means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine and identifying the content of the utterance and the speaker, means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions, means for aggregating the speaking time for each speaker and calculating the percentage of speaking time, means for displaying a text-converted conversation log in real time, means for displaying the emotion analysis results as an emotion icon, means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data, and means for analyzing customer emotions and conversation content during customer service in real time and generating and displaying suggestions for what to say next. This enables store clerks to analyze communication with customers in real time and appropriately determine the optimal response and next action, thereby improving customer satisfaction.

[1305] "Audio data" refers to audio information recorded in digital format.

[1306] "Video data" refers to video information recorded in digital format.

[1307] A "voice recognition engine" is a technology that analyzes voice data and converts it into text data.

[1308] An "facial expression analysis engine" is a technology that analyzes a person's facial expressions from video data and identifies their emotional state.

[1309] "Speech content" is text information obtained by analyzing voice data.

[1310] A "speaker" is a person who speaks each utterance in the voice data.

[1311] "Speech duration" is data indicating the length of time that a particular speaker is speaking.

[1312] The "proportion of speaking time" is the proportion of the time a particular speaker is speaking to the total conversation time.

[1313] A "textualized conversation log" is data that has been recorded by converting voice data into text using a voice recognition engine.

[1314] An "emotion icon" is an icon that visually represents an emotional state identified by the facial expression analysis engine.

[1315] The "evaluation score" is a score obtained by evaluating the quality of speech and responses based on conversation logs and sentiment analysis data.

[1316] "Customer service" refers to the act of a store clerk providing products or services to customers.

[1317] "Customer emotion" refers to the emotional state that a customer displays while being served.

[1318] "Suggestion" is a function that suggests what to talk about next or what to do next.

[1319] The present invention relates to a system that supports customer service in brick-and-mortar stores by acquiring and analyzing voice and video data in real time. In particular, the use of smart glasses enables store clerks to communicate effectively with customers. This system is composed of the following elements: a means for acquiring voice and video data, a voice recognition engine, a facial expression analysis engine, a means for displaying emotional icons, a means for automatically converting conversation logs into text, a means for calculating speech time and evaluation scores, a suggestion function, and a means for accessing a dictionary database.

[1320] Setup and Connection

[1321] The server is connected to the smart glasses and provides tools to assist store clerks in serving customers. When a user (store clerk) starts serving customers, they put on the smart glasses and the system starts automatically, which starts data collection.

[1322] Acquiring audio and video data

[1323] The smart glasses, which are the terminals, use a microphone and camera to capture audio and video data in real time. This data is sent to the server in streaming format. The user simply needs to provide normal customer service without any awareness of the device.

[1324] Speech recognition and talking time percentage calculation

[1325] The server passes the voice data to a speech recognition engine to identify the content and speaker. The speech recognition engine analyzes the data to generate text data and record the content and speaker information. The server then aggregates the speaking time for each speaker and calculates the "talking time ratio" for the user and customer. This information is sent to the device in real time and overlaid on the smart glasses display.

[1326] sentiment analysis

[1327] The server passes the video data to a facial expression analysis engine to identify emotions from facial expressions, and also passes the audio data to a tone of voice analysis engine to identify emotions from tone of voice. These emotion data are integrated and displayed as emotion icons on the user's smart glasses.

[1328] Automatic text conversion of conversation logs

[1329] The server records the text data obtained from the speech recognition engine in real time as a conversation log, which is periodically sent to the terminal and displayed in real time on the user's smart glasses.

[1330] Calculation of evaluation score

[1331] The server evaluates the quality of the conversation and response based on the conversation log and sentiment analysis data, and calculates an evaluation score. This score is displayed in real time on the user's smart glasses, and suggests areas for improvement as needed.

[1332] Suggestion function

[1333] The server analyzes the conversation and progress between the user and the customer, and generates suggestions for next steps and solutions, which are displayed in real time as a pop-up on the user's smart glasses.

[1334] Quick Dictionary Function

[1335] When a user selects a word they don't understand, the device sends the word to the server, which searches for the word's meaning in a dictionary database and sends the results to the user's smart glasses, allowing the user to see the meaning of the selected word in a pop-up window.

[1336] Retrospective training function

[1337] The server stores all data (audio, video, conversation logs, sentiment analysis) during customer service and provides a review interface that users can access later. Based on the stored data, users can replay and analyze past customer service interactions and repeatedly practice with simulation videos that include areas for improvement.

[1338] Specific examples

[1339] For example, if a user serves a customer for 10 minutes, the device sends the audio and video data of the service to the server in real time. The server analyzes the audio data and obtains the results: the user spoke for 6 minutes and the customer spoke for 4 minutes. Based on this, the smart glasses display a speaking time ratio of 60% vs. 40%. At the same time, the content of the conversation is converted into text and displayed in real time on the user's smart glasses as a conversation log. The server also analyzes the customer's facial expressions and tone of voice, and displays emotion icons such as "interested" or "anxious."

[1340] Prompt Sentence Examples

[1341] Here are some example prompts to generate suggestions for the AI:

[1342] Conversation text:

[1343] Salesperson: Welcome. What item are you looking for today?

[1344] Customer: Yes, I'm looking for a gift for Ochugen.

[1345] Suggestion: Do you have a desired price range or a specific genre? You can also take a look at our catalog.

[1346] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1347] Step 1:

[1348] The smart glasses terminal captures audio and video data in real time using a microphone and camera from the moment the customer begins serving customers, and this data is sent to a server in streaming format.

[1349] Input: Real-time audio and video data during customer service

[1350] Output: Streamed audio and video data

[1351] Step 2:

[1352] The server passes the acquired voice data to a voice recognition engine, which analyzes and records the voice data as text data and speaker information. The voice recognition engine uses, for example, Google Speech Recognition.

[1353] Input: Streamed audio data

[1354] Data processing: Converting voice data into text (voice recognition)

[1355] Output: Text data and speaker information

[1356] Step 3:

[1357] At the same time, the server passes the video data to an expression analysis engine (such as DeepFace) to identify emotions from facial expressions, and a voice tone analysis engine to determine emotions from voice.

[1358] Input: Streamed video and audio data

[1359] Data processing: Facial expression analysis from video data, tone of voice analysis from audio data

[1360] Output: Emotion data (facial expressions and tone of voice)

[1361] Step 4:

[1362] The server records the text data obtained from the voice recognition engine as a conversation log, and transmits the textual conversation log to the terminal in real time at specific intervals.

[1363] Input: Text data

[1364] Data processing: Record text data as a conversation log

[1365] Output: Real-time conversation log

[1366] Step 5:

[1367] The server aggregates the speaking time of each speaker and calculates the percentage of time the user (store clerk) is talking and the percentage of time the customer is talking. This information is sent to the device in real time and displayed on the smart glasses display.

[1368] Input: Text data and speaker information

[1369] Data calculation: counting speech time and calculating percentage

[1370] Output: Percentage of speaking time

[1371] Step 6:

[1372] The server combines facial expression data and tone of voice data and sends them to the terminal as an emotional icon, allowing the user to visually grasp the customer's emotions.

[1373] Input: Emotion data (facial expressions and tone of voice)

[1374] Data processing: Emotion data integration and emotion icon generation

[1375] Output:Emotion icon

[1376] Step 7:

[1377] The server calculates an evaluation score based on the conversation log and emotion data and sends it to the device. The evaluation score is calculated based on factors such as the customer's reaction and the balance of their speech.

[1378] Input: Conversation logs and emotion data

[1379] Data calculation: Calculation of evaluation score

[1380] Output: Evaluation score

[1381] Step 8:

[1382] The server analyzes the conversation content and progress and generates suggestions for what to say next. The suggestions are displayed in real time in a pop-up format on the device.

[1383] Input: Conversation log and progress

[1384] Data Computation: Generating Suggestions with Generative AI Models

[1385] Output: Suggestion

[1386] Step 9:

[1387] When a user selects a word they do not understand, the device sends the word to the server, which searches for the meaning of the word in a dictionary database and sends the results to the device, where they are displayed in a pop-up window.

[1388] Input: selected word

[1389] Data processing: Retrieving meanings from dictionary databases

[1390] Output: Word meaning

[1391] Step 10:

[1392] The server stores all data (audio, video, conversation logs, and sentiment analysis data) during customer service and provides a review interface that users can access later. This allows users to replay and analyze past customer service interactions and repeatedly practice areas for improvement using simulation videos.

[1393] Input: Audio data, video data, conversation logs, sentiment analysis data

[1394] Data processing: storing data and providing an interface for reviewing it

[1395] Output: Retrospective interface

[1396] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1397] This invention relates to a multi-functional system that captures and analyzes audio and video data in real time to improve the quality of online business negotiations and interviews. This system provides more efficient communication support by combining it with an emotion engine that recognizes the user's emotions.

[1398] Explanation of program processing

[1399] Setup and Connection

[1400] The server connects the user's device to the main Web meeting product. When the user starts a Web meeting, the device notifies the server of the start of the session. Upon receiving this notification, the server sends the initial setup data for the tool to the user's device, completing the setup of the entire system.

[1401] Acquiring audio and video data

[1402] The device captures audio and video data in real time from its microphone and camera, and sends the captured data in streaming format to the server. The user simply conducts a regular web meeting.

[1403] Speech recognition and talking time percentage calculation

[1404] The server passes the received voice data to a voice recognition engine, which analyzes it to identify the content of the speech and the speaker. The analyzed text data and speaker information are recorded, and the server tallies the speaking time for each speaker and calculates the "percentage of speaking time." This information is sent to the device and displayed as an overlay on the user's screen.

[1405] sentiment analysis

[1406] The server passes video data to a facial expression analysis engine to identify emotions from facial expressions. It also passes audio data to a tone of voice analysis engine to identify emotions from vocal tones. In addition, the present invention incorporates an emotion engine that recognizes the user's emotions and integrates these data to generate an emotion icon. The generated emotion icon is displayed on the user's screen.

[1407] Automatic text conversion of conversation logs

[1408] The server records the conversation content obtained from the speech recognition engine as a text log in real time. The recorded text log is periodically sent to the terminal and displayed on the user's screen in real time.

[1409] Calculation of evaluation score

[1410] The server analyzes the conversation logs and sentiment analysis data to evaluate the user's performance based on best practices. The evaluation score is calculated in real time and sent to the user's device. Based on the evaluation score, the device presents improvements and suggestions to the user.

[1411] Quick Dictionary Function

[1412] When a user selects a word they do not understand, the device sends the word to the server, which searches for the meaning of the word in a dictionary database and sends the results back to the device, allowing the user to see the meaning of the selected word in a pop-up window.

[1413] Suggestion function

[1414] The server analyzes the conversation and its progress, and generates suggestions for what to say next and what questions to ask. These suggestions are sent to the user's device in real time and displayed in a pop-up format.

[1415] Retrospective training function

[1416] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and repeatedly practice areas where they failed or needed improvement using simulation videos.

[1417] Specific examples

[1418] When a user conducts a 20-minute online interview, the device captures audio and video data in real time and sends it to the server. The server analyzes the audio data and calculates the ratio of the user speaking for 12 minutes to the other person speaking for 8 minutes. This information is displayed on the screen. In addition, emotion icons derived from a facial expression analysis engine and a tone of voice analysis engine are displayed on the screen to help the user check for missed questions and consistency in the topic.

[1419] The present invention provides a multi-functional system that further improves the quality of communication by combining it with an emotion engine that recognizes the user's emotions, allowing users to efficiently conduct online business negotiations and online interviews.

[1420] The processing flow will be explained below.

[1421] Explanation of program processing

[1422] Setup and Connection

[1423] Step 1:

[1424] The server connects users' devices and major web meeting products.

[1425] Step 2:

[1426] When a user starts a Web meeting, the terminal notifies the server of the start of the session.

[1427] Step 3:

[1428] The server transmits the initial setting data of the tool to the user's terminal.

[1429] Acquiring audio and video data

[1430] Step 1:

[1431] The device captures audio and video data in real time from the microphone and camera.

[1432] Step 2:

[1433] The terminal transmits the acquired audio and video data to the server in streaming format.

[1434] Speech recognition and talking time percentage calculation

[1435] Step 1:

[1436] The server passes the received voice data to the voice recognition engine.

[1437] Step 2:

[1438] The voice recognition engine analyzes the voice data and identifies what is being said and who is speaking.

[1439] Step 3:

[1440] The server records the analyzed text data and speaker information.

[1441] Step 4:

[1442] The server aggregates the speaking time for each speaker and calculates the "proportion of speaking time" for the user and the other party.

[1443] Step 5:

[1444] The device overlays the percentage of speaking time received from the server on the screen.

[1445] sentiment analysis

[1446] Step 1:

[1447] The server passes the video data to the facial expression analysis engine.

[1448] Step 2:

[1449] The facial expression analysis engine analyzes video data and identifies emotions from facial expressions.

[1450] Step 3:

[1451] The server passes the audio data to a tone of voice analysis engine.

[1452] Step 4:

[1453] A voice tone analysis engine analyzes audio data and identifies emotions from the tone of voice.

[1454] Step 5:

[1455] The server also analyzes the user's facial expressions and voice and passes them to an emotion engine that recognizes the user's emotions.

[1456] Step 6:

[1457] The emotion engine identifies the user's emotion and integrates it with other emotion data.

[1458] Step 7:

[1459] The server generates an emotional icon from the integrated emotional data.

[1460] Step 8:

[1461] The device overlays the emotion icons received from the server on the screen.

[1462] Automatic text conversion of conversation logs

[1463] Step 1:

[1464] The server records the conversation content obtained from the voice recognition engine as a text log in real time.

[1465] Step 2:

[1466] The server periodically sends the text log to the user's terminal.

[1467] Step 3:

[1468] The device overlays the received conversation log on the screen in real time.

[1469] Calculation of evaluation score

[1470] Step 1:

[1471] The server analyzes the conversation logs and sentiment analysis data and evaluates them based on best practice rules.

[1472] Step 2:

[1473] The server calculates the evaluation score and sends the score to the user's terminal.

[1474] Step 3:

[1475] The device displays the received evaluation score on the screen and suggests areas for improvement if necessary.

[1476] Quick Dictionary Function

[1477] Step 1:

[1478] When the user selects a word they do not understand, the terminal sends the word to the server.

[1479] Step 2:

[1480] The server searches the dictionary database for the meaning of the word and sends the result to the terminal.

[1481] Step 3:

[1482] The device will pop up the meaning of the word on the screen.

[1483] Suggestion function

[1484] Step 1:

[1485] The server analyzes the conversation content and progress and generates suggestions for what to say next.

[1486] Step 2:

[1487] The server sends the generated suggestions to the user's terminal.

[1488] Step 3:

[1489] The device displays the received suggestions in a pop-up format on the screen.

[1490] Retrospective training function

[1491] Step 1:

[1492] The server stores all data from the meeting (audio, video, conversation logs, sentiment analysis).

[1493] Step 2:

[1494] The server provides a review interface that users can access after the meeting has ended.

[1495] Step 3:

[1496] Users can access the saved data to replay and analyze past business meetings and interviews.

[1497] Step 4:

[1498] Users can repeatedly practice the areas where they failed or the points they want to improve using simulation videos.

[1499] Specific examples

[1500] For example, if a user conducts a 20-minute online interview, the device captures audio and video data from the microphone and camera in real time and sends it in streaming format to the server. The server analyzes the audio data and calculates the percentage of time the user spoke for 12 minutes and the other party spoke for 8 minutes. This information is overlaid on the user's screen in the form of 68% vs. 32%.

[1501] The server then analyzes the user's facial expressions and tone of voice using an emotion engine, and combines the acquired emotion data to generate an emotion icon. For example, if the server determines that the user is nervous, it will display the emotion icon as "nervous."

[1502] The conversation log is converted into text in real time by a speech recognition engine and displayed on the user's screen. An evaluation score is also calculated in real time, allowing the user to instantly evaluate their own performance.

[1503] When a user uses the quick dictionary feature for a word they don't understand, the meaning of the selected word is displayed in a pop-up, and a suggestion function is also activated to keep the conversation flowing smoothly.

[1504] Finally, after the interview, you can use the retrospective training function to play back past interview data and repeatedly practice areas where you failed or need improvement using simulation footage.

[1505] Example 2

[1506] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1507] In conventional online business meetings and interviews, there was no system that could perform real-time emotion recognition, speech content analysis, evaluation score calculation, next topic suggestion, and even retrospective training all in one place. As a result, users had to rely on multiple tools and manual analysis, making it difficult to improve the quality of efficient communication. In addition, there was little real-time feedback, leading to missed opportunities for improvement.

[1508] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1509] In this invention, the server includes means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine and identifying the content and speaker of the speech, means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions, means for aggregating the speaking time for each speaker and calculating the percentage of speaking time, means for displaying a text-converted conversation log in real time, means for displaying the emotion analysis results as an emotion icon, means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data, means for analyzing the progress of the conversation and generating and displaying suggestions for what to say next, means for searching for and displaying the meaning of a selected word from a dictionary database, and means for unified management of voice, video, conversation log, and emotion analysis data and for replaying and analyzing past business negotiations and interviews based on the stored data. This allows users to receive real-time and comprehensive data analysis and feedback, significantly improving the quality of online business negotiations and online interviews.

[1510] "Audio data" refers to a speaker's audio signal collected through a microphone.

[1511] "Video data" refers to frames of video and their sequences collected through a camera.

[1512] A "voice recognition engine" refers to software or hardware for converting voice data into text data.

[1513] An "facial expression analysis engine" refers to software or hardware that analyzes human facial expressions from video data and identifies their emotions.

[1514] "Utterance content" refers to the specific words uttered by the speaker and their meaning.

[1515] "Speaker" refers to the person speaking.

[1516] "Speech time" refers to the total amount of time a speaker is actually speaking.

[1517] "Proportion of speaking time" refers to the ratio of the time each speaker is speaking to the total time.

[1518] A "conversation log" refers to a record of voice data converted into text.

[1519] "Emotion icons" refers to a display format that shows emotions obtained through facial expression analysis and voice analysis using shapes and emoticons.

[1520] "Evaluation score" refers to the result of quantifying a user's performance based on conversation logs and sentiment analysis data.

[1521] "Suggestions" refer to the next thing to say or questions that are generated based on the progress of the conversation.

[1522] A "dictionary database" refers to a database that associates words with their meanings.

[1523] "Stored Data" refers to audio, video, conversation logs, and sentiment analysis data collected during meetings.

[1524] A "natural language processing engine" refers to software or hardware for analyzing and generating text data.

[1525] This invention relates to a multi-functional system that captures and analyzes audio and video data in real time to improve the quality of online business negotiations and interviews. This system provides more efficient communication support by combining it with an emotion engine that recognizes the user's emotions.

[1526] The main components of this system include the user's device, server, voice recognition engine, facial expression analysis engine, voice tone analysis engine, emotion engine, dictionary database, and natural language processing engine. These elements work together to provide comprehensive meeting support functions.

[1527] The user's device captures audio and video data in real time using a built-in or externally connected microphone and camera. The captured data is sent to the server in a streaming format such as WebRTC. The user simply conducts a regular web meeting, and the data is analyzed by the server.

[1528] The server passes the audio data received in streaming format to a speech recognition engine such as Google Cloud Speech-to-Text for analysis. This identifies the content of the speech and the speaker, and the analyzed text data and speaker information are recorded in a database. The server also tallies the speaking time for each speaker and calculates the "percentage of speaking time." This information is sent to the user's device and displayed on the screen.

[1529] The server then passes the video data to Microsoft Azure's Cognitive Services Facial Expression Recognition API to identify emotions from facial expressions. It also uses IBM Watson's voice tone analysis to identify emotions from voice tone. The emotion engine combines these data to generate emotion icons, which are displayed on the user's screen in real time.

[1530] Furthermore, based on the conversation log and sentiment analysis data, the server evaluates the user's performance based on best practices. This evaluation score is also calculated in real time and sent to the user's device. Based on the evaluation score, the device presents the user with suggestions and areas for improvement.

[1531] The system also includes a quick dictionary function, which means that when a user selects a word they do not understand, the system sends the word to the server, searches for the meaning of the word in the dictionary database, and displays the results in a pop-up.

[1532] The server analyzes the conversation content and progress using a natural language processing (NLP) engine, and generates suggestions for what to say next and what questions to ask. These suggestions are also sent to the user's device in real time and displayed in a pop-up window.

[1533] Even after the meeting, the server stores the audio, video, conversation log, and sentiment analysis data, and provides a review interface that allows users to replay and analyze past business negotiations and interviews based on this data.Users can replay and analyze past business negotiations and interviews based on the saved data, and repeatedly practice areas where they failed or need improvement using simulation videos.

[1534] For example, if a user conducts a 20-minute online interview, the device captures audio and video data in real time and sends it to the server. The server analyzes the audio data and calculates the ratio of the user speaking for 12 minutes to the other person speaking for 8 minutes. This information is displayed on the screen. In addition, emotion icons obtained by a facial expression analysis engine and a tone of voice analysis engine are displayed on the screen to help the user check for missed questions and consistency in the topic.

[1535] An example of a prompt might be, "During the online interview, please tell us how the other person's emotions changed while they were listening to you."

[1536] As described above, the present invention is a system that provides real-time and comprehensive data analysis and feedback in online business negotiations and interviews, enabling users to efficiently improve the quality of their communication.

[1537] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1538] Step 1:

[1539] The device captures audio and video data in real time using a built-in or externally connected microphone and camera. The inputs are audio signals from the microphone and video data from the camera. These data are converted into a streaming format (such as H.264 or AAC) and sent to the server. Specifically, the device driver captures the data and transfers it to the server using the WebRTC protocol.

[1540] Step 2:

[1541] The server passes the voice data received from the device to a voice recognition engine such as Google Cloud Speech-to-Text. The input is streaming voice data. This is sent as an API request, and text data and speaker information are obtained as analysis results. Specifically, the server processes the voice data in batches and sends an API request to the voice recognition engine for each batch. The analyzed text data and speaker information are then recorded in a database.

[1542] Step 3:

[1543] The server passes the received video data to the Microsoft Azure Cognitive Services Facial Expression Recognition API to identify emotions. The input is real-time video data sent from the device. The facial expression analysis engine analyzes the video frames and outputs emotional data. Specifically, it sends an API request for each video frame, maps the obtained emotional data, and records it in a database.

[1544] Step 4:

[1545] The server passes the voice data to IBM Watson's tone of voice analysis engine, which identifies emotions from the tone of voice. The input is streaming voice data. The analysis results in emotional data based on the tone of voice. Specifically, the system sends the voice data in the form of an API request, receives the analyzed emotional data, and records it in a database.

[1546] Step 5:

[1547] The server integrates the emotion data obtained from the video and audio data to generate an emotion icon. The inputs are the results of facial expression analysis and tone of voice analysis. The emotion engine integrates these and outputs an emotion icon. Specifically, the integration algorithm is executed to generate an emotion icon. This icon is then sent to the user's device in real time.

[1548] Step 6:

[1549] The server records the text data obtained from the output of the speech recognition engine as a conversation log and sends it to the device in real time. The input is the analyzed text data. The log recorded in the database is sent to the device via WebSocket at regular intervals and displayed on the user's screen.

[1550] Step 7:

[1551] The server runs an algorithm that evaluates user performance based on conversation logs and sentiment analysis data. The inputs are textual conversation logs and emotional icon data. The server calculates an evaluation score and sends it to the device in real time. Specifically, the server uses a machine learning model to analyze the data and calculate a score.

[1552] Step 8:

[1553] When a user selects a word that the device does not understand, it sends the word to the server and searches for the meaning of the word in the dictionary database. The input is the word selected by the user. The meaning of the word retrieved from the server is displayed in a pop-up. Specifically, it captures the selection operation, sends an API request to the server, retrieves information from the dictionary database, and displays it.

[1554] Step 9:

[1555] The server analyzes the conversation content and progress, and generates the next thing to say and questions. The input is the real-time conversation log and progress data. The natural language processing (NLP) engine performs the analysis, generates suggestions, and sends them to the device. Specifically, the NLP engine analyzes the text data, creates a list of suggestions, and sends them in real time.

[1556] Step 10:

[1557] The server stores all data from the meeting (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. The input is all data collected during the meeting. This is stored in an integrated database so that users can access it later. Specifically, the server stores the data in cloud storage and provides an interface that users can use to play and analyze it through a browser.

[1558] (Application example 2)

[1559] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1560] Today, there are a wide variety of communication methods, but improving the quality of communication is especially important in online business meetings and interviews. However, conventional systems simply collect voice and video data and lack the support to identify emotions and promote effective communication. Furthermore, in situations such as customer support, effective dialogue between support staff and customers is important, and there is a growing need for systems that can analyze and support the progress of conversations in real time.

[1561] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1562] In this invention, the server includes: means for acquiring voice data and video data; means for passing the acquired voice data to a voice recognition engine to identify the content of the utterance and the speaker; means for passing the acquired video data to a facial expression analysis engine to identify emotions from facial expressions; means for aggregating the speaking time for each speaker and calculating the percentage of speaking time; means for displaying a text-converted conversation log in real time; means for displaying the emotion analysis results as emotion icons; means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data; means for analyzing the progress of the conversation and generating and displaying suggestions for what to say next; means for accessing stored data to play back and analyze past contact sessions or interviews; means for searching and displaying the meaning of selected words from a dictionary database; and means for a generative AI model to generate a prompt sentence to say next in real time using the emotion analysis results and the conversation log. This makes it possible to dramatically improve the quality of communication by identifying a user's emotions through real-time analysis of voice and video data and automatically suggesting what to say next.

[1563] "Voice data" is digital information that records the user's speech and surrounding sounds, and is acquired via a voice input device such as a microphone.

[1564] "Video data" is digital information that records visual information, including the movements of objects and people, acquired through a video input device such as a camera.

[1565] A "voice recognition engine" is a technology that analyzes input voice data and converts it into text data.

[1566] An "expression analysis engine" is a technology that analyzes the facial behavior and expressions of people in video data to determine their emotions.

[1567] The term "speaker" is a concept that refers to a specific person speaking in the audio data.

[1568] "Speaking time" refers to the total time that a particular speaker is speaking in the audio data.

[1569] A "textualized conversation log" is a record of the contents of voice data converted into text format.

[1570] An "emotion icon" is an icon that visually represents an emotion identified by the emotion analysis engine.

[1571] The "evaluation score" is a numerical evaluation of a user's performance based on conversation logs and sentiment analysis data.

[1572] "Conversation progress" is the process of analyzing the content of a conversation and its progress in real time.

[1573] "Suggestion" is a function that suggests the next topic or question to talk about based on the progress of the conversation.

[1574] "Stored Data" refers to all audio, video, conversation logs, and sentiment analysis data recorded by the system, including past contact sessions and interviews.

[1575] A "dictionary database" is a database that stores the meanings and related information of words.

[1576] A "generative AI model" is an algorithm generated using machine learning technology that generates the next prompt sentence to be spoken based on the conversation log and sentiment analysis results.

[1577] This invention relates to a multi-functional system that improves the quality of communication in business negotiations and interviews by identifying a user's emotions through real-time analysis of audio data and video data and automatically suggesting what to say next. Specific embodiments of this system will be described below.

[1578] Setup and Connection

[1579] The server connects the user's device to the main web meeting service. When a user starts a web meeting, the device notifies the server of the start of the session. Upon receiving this notification, the server sends the initial setup data for the tool to the user's device, completing the setup of the entire system.

[1580] Acquiring audio and video data

[1581] The device captures audio and video data in real time from its microphone and camera, and sends the captured data in streaming format to the server. The user simply conducts a regular web meeting.

[1582] Speech recognition and speaker identification

[1583] The server passes the received voice data to a voice recognition engine, which analyzes it to identify the content of the speech and the speaker. The analyzed text data and speaker information are recorded, and the server tallies the speaking time for each speaker and calculates the "percentage of speaking time." This information is sent to the device and displayed as an overlay on the user's screen.

[1584] sentiment analysis

[1585] The server passes video data to a facial expression analysis engine to identify emotions from facial expressions. It also passes audio data to a voice tone analysis engine to identify emotions from voice tone. Additionally, the server incorporates an emotion engine that recognizes the user's emotions and combines these data to generate an emotion icon. The generated emotion icon is then displayed on the user's screen.

[1586] Automatic text conversion of conversation logs

[1587] The server records the conversation content obtained from the speech recognition engine as a text log in real time. The recorded text log is periodically sent to the terminal and displayed on the user's screen in real time.

[1588] Calculation of evaluation score

[1589] The server analyzes the conversation logs and sentiment analysis data to evaluate the user's performance based on best practices. The evaluation score is calculated in real time and sent to the user's device. Based on the evaluation score, the device presents improvements and suggestions to the user.

[1590] Quick Dictionary Function

[1591] When a user selects a word they do not understand, the device sends the word to the server, which searches for the meaning of the word in a dictionary database and sends the results back to the device, allowing the user to see the meaning of the selected word in a pop-up window.

[1592] Suggestion function

[1593] The server analyzes the conversation content and progress, and generates suggestions for what to say next or what questions to ask. These suggestions are sent to the user's device in real time and displayed in a pop-up format. At this time, a generative AI model can be used to generate prompts. For example, a prompt such as "What should be the next question after this topic?" can be generated.

[1594] Retrospective training function

[1595] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and repeatedly practice areas where they failed or needed improvement using simulation videos.

[1596] Specific examples

[1597] A customer support agent conducts a 20-minute web chat session, during which the system analyzes audio and video data to identify agent and customer sentiment. Based on the conversation's progress, the system displays appropriate next questions and suggestions in real time, and provides data for retrospective training after the session ends. Using a generative AI model, the system can generate prompts such as, "What's the next question we should ask about this topic?"

[1598] As described above, this system can dramatically improve the quality of communication by identifying the user's emotions through real-time analysis of voice and video data and automatically suggesting what to say next.

[1599] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1600] Step 1: Setup and Connection

[1601] The server connects the user's terminal to the main Web meeting service. When a user starts a Web meeting, the terminal notifies the server of the start of the session. The input is the session start notification data, and the output is the initial setup data for the tool. The server receives this notification and sends the initial setup data for the tool to the user's terminal, completing the setup of the entire system.

[1602] Step 2: Acquiring audio and video data

[1603] The device captures audio and video data in real time from the microphone and camera. The captured data is sent to the server in streaming format as input data. The output is streaming audio and video data. The user simply needs to proceed with a normal web meeting.

[1604] Step 3: Speech recognition and speaker identification

[1605] The server passes the received voice data to a voice recognition engine for analysis. This analysis identifies the content of the speech and the speaker. The input is stream-format voice data, and the output is text data and speaker information. The server records this data and tallys up the speaking time for each speaker.

[1606] Step 4: Sentiment analysis

[1607] The server passes video data to a facial expression analysis engine to identify emotions from facial expressions. It also passes audio data to a tone of voice analysis engine to identify emotions from vocal tone. The input is streamed video and audio data, and the output is the emotion analysis results. These data are integrated to generate an emotional icon. The generated emotional icon is sent to the terminal and displayed on the user's screen.

[1608] Step 5: Automatic transcription of conversation logs

[1609] The server records the conversation content obtained from the speech recognition engine as a text log in real time. The input is text data and the output is a text log. This text log is periodically sent to the terminal and displayed on the user's screen in real time.

[1610] Step 6: Calculating the evaluation score

[1611] The server analyzes the conversation log and sentiment analysis data and evaluates the user's performance based on best practices. The input is the conversation log and sentiment analysis data, and the output is an evaluation score. The evaluation score is calculated in real time and sent to the user's device. The device then presents the user with suggestions and areas for improvement based on the evaluation score.

[1612] Step 7: Quick Dictionary Function

[1613] When the user selects a word that the terminal does not understand, it sends the word to the server. The input is the selected word, and the output is the meaning of the word. The server searches for the meaning of the word in the dictionary database and sends the result to the terminal. The user is prompted to confirm the meaning of the selected word in a pop-up window.

[1614] Step 8: Suggestions

[1615] The server analyzes the conversation content and progress, and generates suggestions for what to say next and what questions to ask. At this time, a generative AI model can be used to generate prompts. The input is the conversation log and progress, and the output is suggestions and prompts. These suggestions are sent to the device in real time and displayed in a pop-up format.

[1616] Step 9: Retrospective Training Function

[1617] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. The input is the meeting data, and the output is the review interface and past data. Based on the stored data, users can replay and analyze past business negotiations and interviews, and repeatedly practice using simulation videos to identify areas where they failed or need improvement.

[1618] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1619] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1620] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1621] [Fourth embodiment]

[1622] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1623] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1624] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1625] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1626] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1627] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1628] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1629] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1630] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1631] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1632] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1633] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1634] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1635] This invention relates to a system that supports online business negotiations and online interviews by acquiring and analyzing voice and video data in real time. This system is mainly composed of the following elements: a means for acquiring voice and video data, a voice recognition engine, a facial expression analysis engine, a means for displaying emotional icons, a means for automatically converting conversation logs into text, a means for calculating speech time and evaluation scores, a suggestion function, and a means for accessing a dictionary database.

[1636] Explanation of program processing

[1637] Setup and Connection

[1638] The server connects the user's device to a major web meeting product (such as a general web conferencing system). When the user starts a web meeting, the device connects to the server and launches a support tool, which starts collecting data for business negotiations or interviews.

[1639] Acquiring audio and video data

[1640] The device captures audio and video data in real time from the microphone and camera. This data is then sent directly to the server in streaming format. The user is then able to proceed with a normal web meeting without being aware of anything.

[1641] Speech recognition and talking time percentage calculation

[1642] The server passes the voice data to a voice recognition engine, which identifies the content of the speech and the speaker. The voice recognition engine analyzes the data, generates text data, and records the content of the speech and speaker information. The server then aggregates the speaking time for each speaker and calculates the "talking time ratio" for the user and the other party. This information is sent to the device in real time and displayed as an overlay on the screen.

[1643] sentiment analysis

[1644] The server passes the video data to a facial expression analysis engine, which identifies emotions from facial expressions. The audio data is also passed to a tone of voice analysis engine, which identifies emotions from the tone of voice. These emotion data are integrated and displayed on the user's screen as an emotion icon.

[1645] Automatic text conversion of conversation logs

[1646] The server records the text data obtained from the speech recognition engine in real time as a conversation log, which is periodically sent to the terminal and displayed on the user's screen in real time.

[1647] Calculation of evaluation score

[1648] The server evaluates the conversation logs and sentiment analysis data based on best practice rules to calculate a rating score, which is displayed in real time on the user's device and suggests areas for improvement if necessary.

[1649] Quick Dictionary Function

[1650] When a user selects a word they don't understand, the device sends it to the server, which searches for the word's meaning in the dictionary database and sends the results to the user's device, where the user can see the meaning of the selected word in a pop-up window.

[1651] Suggestion function

[1652] The server analyzes the conversation and its progress, and generates suggestions for what to say next and what questions to ask. These suggestions are sent to the user's device in real time and displayed in a pop-up window.

[1653] Retrospective training function

[1654] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and practice repeatedly using simulation videos that include areas of failure and areas for improvement.

[1655] Specific examples

[1656] For example, consider a case where a user has a 10-minute online business meeting. The device sends the audio and video data of the meeting to the server in real time. The server analyzes the audio data and obtains the results: the user spoke for 6 minutes and the other party spoke for 4 minutes. Based on this, the speaking time ratios of 68% vs. 32% are displayed on the screen. At the same time, the content of the conversation is converted into text and displayed in real time on the user's screen as a conversation log. The server also analyzes the other party's facial expressions and tone of voice, and displays emotion icons such as "skepticism" or "interest."

[1657] In this way, the present invention is a multifunctional, real-time support system that enables users to conduct online business negotiations and online interviews more effectively.

[1658] The processing flow will be explained below.

[1659] Explanation of program processing

[1660] Setup and Connection

[1661] Step 1:

[1662] The server connects users' devices and major web meeting products.

[1663] Step 2:

[1664] When a user starts a Web meeting, the terminal notifies the server of the start of the session.

[1665] Step 3:

[1666] The server transmits the initial setting data of the tool to the user's terminal.

[1667] Acquiring audio and video data

[1668] Step 1:

[1669] The device captures audio and video data in real time from the microphone and camera.

[1670] Step 2:

[1671] The terminal transmits the acquired audio and video data to the server in streaming format.

[1672] Speech recognition and talking time percentage calculation

[1673] Step 1:

[1674] The server passes the received voice data to the voice recognition engine.

[1675] Step 2:

[1676] The voice recognition engine analyzes the voice data and identifies what is being said and who is speaking.

[1677] Step 3:

[1678] The server records the analyzed text data and speaker information.

[1679] Step 4:

[1680] The server aggregates the speaking time for each speaker and calculates the "proportion of speaking time" for the user and the other party.

[1681] Step 5:

[1682] The device overlays the percentage of speaking time received from the server on the screen.

[1683] sentiment analysis

[1684] Step 1:

[1685] The server passes the video data to the facial expression analysis engine.

[1686] Step 2:

[1687] The facial expression analysis engine analyzes video data and identifies emotions from facial expressions.

[1688] Step 3:

[1689] The server passes the audio data to a tone of voice analysis engine.

[1690] Step 4:

[1691] A voice tone analysis engine analyzes audio data and identifies emotions from the tone of voice.

[1692] Step 5:

[1693] The server integrates the obtained emotion data and generates an emotion icon.

[1694] Step 6:

[1695] The device overlays the emotion icons received from the server on the screen.

[1696] Automatic text conversion of conversation logs

[1697] Step 1:

[1698] The server records the conversation content obtained from the voice recognition engine as a text log in real time.

[1699] Step 2:

[1700] The server periodically sends the text log to the user's terminal.

[1701] Step 3:

[1702] The device overlays the received conversation log on the screen in real time.

[1703] Calculation of evaluation score

[1704] Step 1:

[1705] The server analyzes the conversation logs and sentiment analysis data and evaluates them based on best practice rules.

[1706] Step 2:

[1707] The server calculates the evaluation score and sends the score to the user's terminal.

[1708] Step 3:

[1709] The device displays the received evaluation score on the screen and suggests areas for improvement if necessary.

[1710] Quick Dictionary Function

[1711] Step 1:

[1712] When the user selects a word they do not understand, the terminal sends the word to the server.

[1713] Step 2:

[1714] The server searches the dictionary database for the meaning of the word and sends the result to the terminal.

[1715] Step 3:

[1716] The device will pop up the meaning of the word on the screen.

[1717] Suggestion function

[1718] Step 1:

[1719] The server analyzes the conversation content and progress and generates suggestions for what to say next.

[1720] Step 2:

[1721] The server sends the generated suggestions to the user's terminal.

[1722] Step 3:

[1723] The device displays the received suggestions in a pop-up format on the screen.

[1724] Retrospective training function

[1725] Step 1:

[1726] The server stores all data from the meeting (audio, video, conversation logs, sentiment analysis).

[1727] Step 2:

[1728] The server provides a review interface that users can access after the meeting has ended.

[1729] Step 3:

[1730] Users can access the saved data to replay and analyze past business meetings and interviews.

[1731] Step 4:

[1732] Users can repeatedly practice the areas where they failed or the points they want to improve using simulation videos.

[1733] Example 1

[1734] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1735] In recent years, remote work and online business negotiations and interviews have become more common, but communication in these online environments is considered more difficult than face-to-face communication. It is particularly difficult to grasp the other person's reactions and emotions in real time, making it difficult to effectively progress and evaluate the dialogue. Therefore, there is a demand for a system that supports online business negotiations and interviews, with a function that can analyze the other person's reactions and emotions in real time and enable users to communicate effectively.

[1736] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1737] In this invention, the server includes means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine and identifying the content of the speech and the speaker, means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions, means for aggregating the speaking time for each speaker and calculating the percentage of speaking time, means for displaying a textualized conversation log in real time, means for displaying the emotion analysis results as emotion icons, means for recording the text data acquired from the voice recognition engine as a conversation log and transmitting it to a terminal, and means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data. This makes it possible to grasp the other party's reactions and emotions in real time and provide an environment in which users can communicate effectively online.

[1738] "Audio data" refers to sound information acquired through a microphone during online business negotiations or online interviews.

[1739] "Video data" refers to video information captured through a camera during online business negotiations or interviews.

[1740] A "voice recognition engine" is a general term for technology that automatically identifies the content and speaker of an utterance from acquired voice data and converts it into text data.

[1741] "Facial expression analysis engine" is a general term for technology that analyzes facial expressions from acquired video data and identifies emotions from them.

[1742] "Speech duration" refers to the length of time during which a particular speaker speaks.

[1743] "Text conversion" is the process of analyzing audio data and recording what is being said as written information.

[1744] A "conversation log" is a record of text data generated by a voice recognition engine, recorded in chronological order.

[1745] "Emotion icons" are shapes or facial expression icons that visually represent the results of the emotion analysis engine.

[1746] The "evaluation score" is a numerical evaluation value that represents the quality and effectiveness of a conversation based on conversation logs and sentiment analysis data.

[1747] "Suggestion" is a function that analyzes the content and progress of a conversation and suggests what to talk about next or what questions to ask.

[1748] A "dictionary database" is a database that stores the meanings and explanations of words.

[1749] "Real-time display" refers to the act of displaying acquired data and analysis results instantly with almost no delay.

[1750] "Cloud storage" is an online storage service for storing data over the Internet.

[1751] This invention is a system that supports online business negotiations and online interviews by capturing and analyzing audio and video data in real time. This system has the following configuration, including the following main hardware and software elements: a microphone, a camera, a voice recognition engine, a facial expression analysis engine, a suggestion function, and a dictionary database access means.

[1752] Setup and Connection

[1753] The server connects the user's device to the main web meeting product. When the user starts a web meeting, the device connects to the server and launches the support tool, which starts collecting data for the business meeting or interview.

[1754] Acquiring audio and video data

[1755] The device uses a microphone and camera to capture audio and video data in real time, which is then sent in streaming format to a server for analysis.

[1756] Speech recognition and talking time percentage calculation

[1757] The server passes the voice data to a voice recognition engine, which identifies the content of the speech and the speaker. The voice recognition engine converts the voice data into text data and records the content of the speech and speaker information. The server then aggregates the speaking time for each speaker and calculates the "talking time ratio" for the user and the other party. This information is sent to the device in real time and displayed as an overlay on the screen.

[1758] As a specific example of use, the voice recognition engine uses the Google Cloud Speech-to-Text API to convert actual voice data into text with extremely high accuracy.

[1759] sentiment analysis

[1760] The server passes the video data to a facial expression analysis engine, which identifies emotions from facial expressions. The audio data is passed to a tone of voice analysis engine, which identifies emotions from the tone of voice. These emotion data are integrated and displayed on the user's screen as an emotion icon.

[1761] As a specific example of use, the facial expression analysis engine uses Microsoft Azure Face API, and the voice tone analysis engine uses IBM Watson Tone Analyzer.

[1762] Automatic text conversion of conversation logs

[1763] The server records the text data obtained from the speech recognition engine in real time as a conversation log, which is periodically sent to the terminal and displayed on the user's screen in real time.

[1764] Calculation of evaluation score

[1765] The server evaluates the conversation log and sentiment analysis data using best practice rules to calculate a score, which is displayed on the device in real time and suggests areas for improvement if necessary.

[1766] Quick Dictionary Function

[1767] When a user selects a word they do not understand, the device sends it to the server, which searches for the meaning of the word in a dictionary database and sends the results to the device, allowing the user to check the meaning of the selected word in a pop-up window.

[1768] Suggestion function

[1769] The server analyzes the conversation and progress, and generates suggestions for what to say next and what questions to ask. These suggestions are sent to the device in real time and displayed in a pop-up format.

[1770] Retrospective training function

[1771] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and practice repeatedly using simulation videos that include areas of failure and areas for improvement.

[1772] Specific examples

[1773] For example, if a user has a 10-minute online business meeting, the device sends the audio and video data of the meeting to the server in real time. The server analyzes the audio data and determines that the user spoke for 6 minutes and the other party spoke for 4 minutes. Based on this, the speaking time ratio of 68% vs. 32% is displayed on the screen. At the same time, the content of the conversation is converted into text and displayed on the user's screen in real time as a conversation log. The server also analyzes the other party's facial expressions and tone of voice, and displays emotion icons such as "suspicious" and "interested."

[1774] Prompt Sentence Examples

[1775] "Please explain about a system that allows users to check emotional icons and conversation logs in real time during online business negotiations and suggests ways to improve their conversation skills based on the other party's reactions."

[1776] In this way, the present invention is a multi-functional, real-time support system that enables users to conduct online business negotiations and online interviews more effectively.

[1777] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1778] Step 1: Setup and Connection

[1779] The server connects the user's device to the main Web meeting product. When the user starts a Web meeting, the device works with the server to launch the support tool.

[1780] Specific behavior:

[1781] Input: User action to start a web meeting

[1782] Data processing: The server connects to the meeting using the API of the web conferencing system.

[1783] Output: Gets meeting session information and sends it to the device

[1784] Step 2: Acquiring audio and video data

[1785] The device uses a microphone and camera to capture audio and video data in real time, and transmits this data to a server in streaming format.

[1786] Specific behavior:

[1787] Input: Audio data (from microphone), video data (from camera)

[1788] Data Processing: Real-time acquisition and synchronization of audio and video data

[1789] Output: Sends audio and video data to the server in streaming format

[1790] Step 3: Speech recognition and speaking time calculation

[1791] The server passes the voice data to a speech recognition engine to identify the content and speaker, then aggregates the speaking time for each speaker and calculates the "speaking time percentage."

[1792] Specific behavior:

[1793] Input: A stream of audio data

[1794] Data calculation: Converting voice data into text using the Google Cloud Speech-to-Text API, and identifying the content and speaker

[1795] Output: Send the aggregated results to the terminal and overlay them on the screen

[1796] Step 4: Sentiment analysis

[1797] The server passes the video data to an expression analysis engine to determine emotions from facial expressions, and passes the audio data to a tone of voice analysis engine to determine emotions.

[1798] Specific behavior:

[1799] Input: Video data (to facial expression analysis engine), audio data (to voice tone analysis engine)

[1800] Data calculation: Facial expression analysis using Microsoft Azure Face API, voice tone analysis using IBM Watson Tone Analyzer

[1801] Output: Generate emoticons and send them to the device

[1802] Step 5: Automatic transcription of conversation logs

[1803] The server records the text data obtained from the voice recognition engine as a conversation log and periodically sends it to the terminal.

[1804] Specific behavior:

[1805] Input: Text data from the speech recognition engine

[1806] Data processing: Recorded in a database as a conversation log

[1807] Output: Real-time updated conversation log sent to device

[1808] Step 6: Calculating the evaluation score

[1809] The server calculates a rating score using best practice rules based on conversation logs and sentiment analysis data.

[1810] Specific behavior:

[1811] Input: Conversation logs, sentiment analysis data

[1812] Data calculations: evaluated based on best practice rules

[1813] Output: Send the evaluation score to the terminal and display it.

[1814] Step 7: Quick Dictionary Function

[1815] The terminal sends the word selected by the user to the server, which searches for the meaning of the word from a dictionary database and returns it.

[1816] Specific behavior:

[1817] Input: Selected word (from terminal)

[1818] Data operations: Retrieving word meanings from dictionary databases (e.g., Merriam-Webster API)

[1819] Output: Send word meanings to terminal and display

[1820] Step 8: Suggestions

[1821] Based on the analysis of the conversation content and the progress, the server generates the next topic to talk about and questions, and sends them to the user's device in real time.

[1822] Specific behavior:

[1823] Input: Conversation content and progress data

[1824] Data Computation: Using generative AI models to generate suggestions

[1825] Output: Send suggestions to the device and display them in a popup

[1826] Step 9: Retrospective Training Function

[1827] The server stores all data from the meeting and allows users to access a review interface after the meeting, allowing them to replay and analyze past data and identify areas for improvement.

[1828] Specific behavior:

[1829] Input: Meeting audio, video, conversation logs, and sentiment analysis data

[1830] Data processing: Save data to cloud storage

[1831] Output: Provides a retrospective interface that allows users to play back and analyze data

[1832] (Application example 1)

[1833] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1834] Conventional customer service systems lack real-time feedback and support to enable store clerks to communicate smoothly with customers. This can lead to an inability to respond quickly to customer needs, which can lead to a decline in customer satisfaction. This can also have a negative impact on store sales and repeat customer rates. The present invention aims to solve these problems and provide a system that enables store clerks to communicate more effectively and smoothly with customers.

[1835] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1836] In this invention, the server includes means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine and identifying the content of the utterance and the speaker, means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions, means for aggregating the speaking time for each speaker and calculating the percentage of speaking time, means for displaying a text-converted conversation log in real time, means for displaying the emotion analysis results as an emotion icon, means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data, and means for analyzing customer emotions and conversation content during customer service in real time and generating and displaying suggestions for what to say next. This enables store clerks to analyze communication with customers in real time and appropriately determine the optimal response and next action, thereby improving customer satisfaction.

[1837] "Audio data" refers to audio information recorded in digital format.

[1838] "Video data" refers to video information recorded in digital format.

[1839] A "voice recognition engine" is a technology that analyzes voice data and converts it into text data.

[1840] An "facial expression analysis engine" is a technology that analyzes a person's facial expressions from video data and identifies their emotional state.

[1841] "Speech content" is text information obtained by analyzing voice data.

[1842] A "speaker" is a person who speaks each utterance in the voice data.

[1843] "Speech duration" is data indicating the length of time that a particular speaker is speaking.

[1844] The "proportion of speaking time" is the proportion of the time a particular speaker is speaking to the total conversation time.

[1845] A "textualized conversation log" is data that has been recorded by converting voice data into text using a voice recognition engine.

[1846] An "emotion icon" is an icon that visually represents an emotional state identified by the facial expression analysis engine.

[1847] The "evaluation score" is a score obtained by evaluating the quality of speech and responses based on conversation logs and sentiment analysis data.

[1848] "Customer service" refers to the act of a store clerk providing products or services to customers.

[1849] "Customer emotion" refers to the emotional state that a customer displays while being served.

[1850] "Suggestion" is a function that suggests what to talk about next or what to do next.

[1851] The present invention relates to a system that supports customer service in brick-and-mortar stores by acquiring and analyzing voice and video data in real time. In particular, the use of smart glasses enables store clerks to communicate effectively with customers. This system is composed of the following elements: a means for acquiring voice and video data, a voice recognition engine, a facial expression analysis engine, a means for displaying emotional icons, a means for automatically converting conversation logs into text, a means for calculating speech time and evaluation scores, a suggestion function, and a means for accessing a dictionary database.

[1852] Setup and Connection

[1853] The server is connected to the smart glasses and provides tools to assist store clerks in serving customers. When a user (store clerk) starts serving customers, they put on the smart glasses and the system starts automatically, which starts data collection.

[1854] Acquiring audio and video data

[1855] The smart glasses, which are the terminals, use a microphone and camera to capture audio and video data in real time. This data is sent to the server in streaming format. The user simply needs to provide normal customer service without any awareness of the device.

[1856] Speech recognition and talking time percentage calculation

[1857] The server passes the voice data to a speech recognition engine to identify the content and speaker. The speech recognition engine analyzes the data to generate text data and record the content and speaker information. The server then aggregates the speaking time for each speaker and calculates the "talking time ratio" for the user and customer. This information is sent to the device in real time and overlaid on the smart glasses display.

[1858] sentiment analysis

[1859] The server passes the video data to a facial expression analysis engine to identify emotions from facial expressions, and also passes the audio data to a tone of voice analysis engine to identify emotions from tone of voice. These emotion data are integrated and displayed as emotion icons on the user's smart glasses.

[1860] Automatic text conversion of conversation logs

[1861] The server records the text data obtained from the speech recognition engine in real time as a conversation log, which is periodically sent to the terminal and displayed in real time on the user's smart glasses.

[1862] Calculation of evaluation score

[1863] The server evaluates the quality of the conversation and response based on the conversation log and sentiment analysis data, and calculates an evaluation score. This score is displayed in real time on the user's smart glasses, and suggests areas for improvement as needed.

[1864] Suggestion function

[1865] The server analyzes the conversation and progress between the user and the customer, and generates suggestions for next steps and solutions, which are displayed in real time as a pop-up on the user's smart glasses.

[1866] Quick Dictionary Function

[1867] When a user selects a word they don't understand, the device sends the word to the server, which searches for the word's meaning in a dictionary database and sends the results to the user's smart glasses, allowing the user to see the meaning of the selected word in a pop-up window.

[1868] Retrospective training function

[1869] The server stores all data (audio, video, conversation logs, sentiment analysis) during customer service and provides a review interface that users can access later. Based on the stored data, users can replay and analyze past customer service interactions and repeatedly practice with simulation videos that include areas for improvement.

[1870] Specific examples

[1871] For example, if a user serves a customer for 10 minutes, the device sends the audio and video data of the service to the server in real time. The server analyzes the audio data and obtains the results: the user spoke for 6 minutes and the customer spoke for 4 minutes. Based on this, the smart glasses display a speaking time ratio of 60% vs. 40%. At the same time, the content of the conversation is converted into text and displayed in real time on the user's smart glasses as a conversation log. The server also analyzes the customer's facial expressions and tone of voice, and displays emotion icons such as "interested" or "anxious."

[1872] Prompt Sentence Examples

[1873] Here are some example prompts to generate suggestions for the AI:

[1874] Conversation text:

[1875] Salesperson: Welcome. What item are you looking for today?

[1876] Customer: Yes, I'm looking for a gift for Ochugen.

[1877] Suggestion: Do you have a desired price range or a specific genre? You can also take a look at our catalog.

[1878] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1879] Step 1:

[1880] The smart glasses terminal captures audio and video data in real time using a microphone and camera from the moment the customer begins serving customers, and this data is sent to a server in streaming format.

[1881] Input: Real-time audio and video data during customer service

[1882] Output: Streamed audio and video data

[1883] Step 2:

[1884] The server passes the acquired voice data to a voice recognition engine, which analyzes and records the voice data as text data and speaker information. The voice recognition engine uses, for example, Google Speech Recognition.

[1885] Input: Streamed audio data

[1886] Data processing: Converting voice data into text (voice recognition)

[1887] Output: Text data and speaker information

[1888] Step 3:

[1889] At the same time, the server passes the video data to an expression analysis engine (such as DeepFace) to identify emotions from facial expressions, and a voice tone analysis engine to determine emotions from voice.

[1890] Input: Streamed video and audio data

[1891] Data processing: Facial expression analysis from video data, tone of voice analysis from audio data

[1892] Output: Emotion data (facial expressions and tone of voice)

[1893] Step 4:

[1894] The server records the text data obtained from the voice recognition engine as a conversation log, and transmits the textual conversation log to the terminal in real time at specific intervals.

[1895] Input: Text data

[1896] Data processing: Record text data as a conversation log

[1897] Output: Real-time conversation log

[1898] Step 5:

[1899] The server aggregates the speaking time of each speaker and calculates the percentage of time the user (store clerk) is talking and the percentage of time the customer is talking. This information is sent to the device in real time and displayed on the smart glasses display.

[1900] Input: Text data and speaker information

[1901] Data calculation: counting speech time and calculating percentage

[1902] Output: Percentage of speaking time

[1903] Step 6:

[1904] The server combines facial expression data and tone of voice data and sends them to the terminal as an emotional icon, allowing the user to visually grasp the customer's emotions.

[1905] Input: Emotion data (facial expressions and tone of voice)

[1906] Data processing: Emotion data integration and emotion icon generation

[1907] Output:Emotion icon

[1908] Step 7:

[1909] The server calculates an evaluation score based on the conversation log and emotion data and sends it to the device. The evaluation score is calculated based on factors such as the customer's reaction and the balance of their speech.

[1910] Input: Conversation logs and emotion data

[1911] Data calculation: Calculation of evaluation score

[1912] Output: Evaluation score

[1913] Step 8:

[1914] The server analyzes the conversation content and progress and generates suggestions for what to say next. The suggestions are displayed in real time in a pop-up format on the device.

[1915] Input: Conversation log and progress

[1916] Data Computation: Generating Suggestions with Generative AI Models

[1917] Output: Suggestion

[1918] Step 9:

[1919] When a user selects a word they do not understand, the device sends the word to the server, which searches for the meaning of the word in a dictionary database and sends the results to the device, where they are displayed in a pop-up window.

[1920] Input: selected word

[1921] Data processing: Retrieving meanings from dictionary databases

[1922] Output: Word meaning

[1923] Step 10:

[1924] The server stores all data (audio, video, conversation logs, and sentiment analysis data) during customer service and provides a review interface that users can access later. This allows users to replay and analyze past customer service interactions and repeatedly practice areas for improvement using simulation videos.

[1925] Input: Audio data, video data, conversation logs, sentiment analysis data

[1926] Data processing: storing data and providing an interface for reviewing it

[1927] Output: Retrospective interface

[1928] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1929] This invention relates to a multi-functional system that captures and analyzes audio and video data in real time to improve the quality of online business negotiations and interviews. This system provides more efficient communication support by combining it with an emotion engine that recognizes the user's emotions.

[1930] Explanation of program processing

[1931] Setup and Connection

[1932] The server connects the user's device to the main Web meeting product. When the user starts a Web meeting, the device notifies the server of the start of the session. Upon receiving this notification, the server sends the initial setup data for the tool to the user's device, completing the setup of the entire system.

[1933] Acquiring audio and video data

[1934] The device captures audio and video data in real time from its microphone and camera, and sends the captured data in streaming format to the server. The user simply conducts a regular web meeting.

[1935] Speech recognition and talking time percentage calculation

[1936] The server passes the received voice data to a voice recognition engine, which analyzes it to identify the content of the speech and the speaker. The analyzed text data and speaker information are recorded, and the server tallies the speaking time for each speaker and calculates the "percentage of speaking time." This information is sent to the device and displayed as an overlay on the user's screen.

[1937] sentiment analysis

[1938] The server passes video data to a facial expression analysis engine to identify emotions from facial expressions. It also passes audio data to a tone of voice analysis engine to identify emotions from vocal tones. In addition, the present invention incorporates an emotion engine that recognizes the user's emotions and integrates these data to generate an emotion icon. The generated emotion icon is displayed on the user's screen.

[1939] Automatic text conversion of conversation logs

[1940] The server records the conversation content obtained from the speech recognition engine as a text log in real time. The recorded text log is periodically sent to the terminal and displayed on the user's screen in real time.

[1941] Calculation of evaluation score

[1942] The server analyzes the conversation logs and sentiment analysis data to evaluate the user's performance based on best practices. The evaluation score is calculated in real time and sent to the user's device. Based on the evaluation score, the device presents improvements and suggestions to the user.

[1943] Quick Dictionary Function

[1944] When a user selects a word they do not understand, the device sends the word to the server, which searches for the meaning of the word in a dictionary database and sends the results back to the device, allowing the user to see the meaning of the selected word in a pop-up window.

[1945] Suggestion function

[1946] The server analyzes the conversation and its progress, and generates suggestions for what to say next and what questions to ask. These suggestions are sent to the user's device in real time and displayed in a pop-up format.

[1947] Retrospective training function

[1948] The server stores all meeting data (audio, video, conversation logs, sentiment analysis) and provides a review interface that users can access after the meeting. Based on the stored data, users can replay and analyze past business negotiations and interviews, and repeatedly practice areas where they failed or needed improvement using simulation videos.

[1949] Specific examples

[1950] When a user conducts a 20-minute online interview, the device captures audio and video data in real time and sends it to the server. The server analyzes the audio data and calculates the ratio of the user speaking for 12 minutes to the other person speaking for 8 minutes. This information is displayed on the screen. In addition, emotion icons derived from a facial expression analysis engine and a tone of voice analysis engine are displayed on the screen to help the user check for missed questions and consistency in the topic.

[1951] The present invention provides a multi-functional system that further improves the quality of communication by combining it with an emotion engine that recognizes the user's emotions, allowing users to efficiently conduct online business negotiations and online interviews.

[1952] The processing flow will be explained below.

[1953] Explanation of program processing

[1954] Setup and Connection

[1955] Step 1:

[1956] The server connects users' devices and major web meeting products.

[1957] Step 2:

[1958] When a user starts a Web meeting, the terminal notifies the server of the start of the session.

[1959] Step 3:

[1960] The server transmits the initial setting data of the tool to the user's terminal.

[1961] Acquiring audio and video data

[1962] Step 1:

[1963] The device captures audio and video data in real time from the microphone and camera.

[1964] Step 2:

[1965] The terminal transmits the acquired audio and video data to the server in streaming format.

[1966] Speech recognition and talking time percentage calculation

[1967] Step 1:

[1968] The server passes the received voice data to the voice recognition engine.

[1969] Step 2:

[1970] The voice recognition engine analyzes the voice data and identifies what is being said and who is speaking.

[1971] Step 3:

[1972] The server records the analyzed text data and speaker information.

[1973] Step 4:

[1974] The server aggregates the speaking time for each speaker and calculates the "proportion of speaking time" for the user and the other party.

[1975] Step 5:

[1976] The device overlays the percentage of speaking time received from the server on the screen.

[1977] sentiment analysis

[1978] Step 1:

[1979] The server passes the video data to the facial expression analysis engine.

[1980] Step 2:

[1981] The facial expression analysis engine analyzes video data and identifies emotions from facial expressions.

[1982] Step 3:

[1983] The server passes the audio data to a tone of voice analysis engine.

[1984] Step 4:

[1985] A voice tone analysis engine analyzes audio data and identifies emotions from the tone of voice.

[1986] Step 5:

[1987] The server also analyzes the user's facial expressions and voice and passes them to an emotion engine that recognizes the user's emotions.

[1988] Step 6:

[1989] The emotion engine identifies the user's emotion and integrates it with other emotion data.

[1990] Step 7:

[1991] The server generates an emotional icon from the integrated emotional data.

[1992] Step 8:

[1993] The device overlays the emotion icons received from the server on the screen.

[1994] Automatic text conversion of conversation logs

[1995] Step 1:

[1996] The server records the conversation content obtained from the voice recognition engine as a text log in real time.

[1997] Step 2:

[1998] The server periodically sends the text log to the user's terminal.

[1999] Step 3:

[2000] The device overlays the received conversation log on the screen in real time.

[2001] Calculation of evaluation score

[2002] Step 1:

[2003] The server analyzes the conversation logs and sentiment analysis data and evaluates them based on best practice rules.

[2004] Step 2:

[2005] The server calculates the evaluation score and sends the score to the user's terminal.

[2006] Step 3:

[2007] The device displays the received evaluation score on the screen and suggests areas for improvement if necessary.

[2008] Quick Dictionary Function

[2009] Step 1:

[2010] When the user selects a word they do not understand, the terminal sends the word to the server.

[2011] Step 2:

[2012] The server searches the dictionary database for the meaning of the word and sends the result to the terminal.

[2013] Step 3:

[2014] The device will pop up the meaning of the word on the screen.

[2015] Suggestion function

[2016] Step 1:

[2017] The server analyzes the conversation content and progress and generates suggestions for what to say next.

[2018] Step 2:

[2019] The server sends the generated suggestions to the user's terminal.

[2020] Step 3:

[2021] The device displays the received suggestions in a pop-up format on the screen.

[2022] Retrospective training function

[2023] Step 1:

[2024] The server stores all data from the meeting (audio, video, conversation logs, sentiment analysis).

[2025] Step 2:

[2026] The server provides a review interface that users can access after the meeting has ended.

[2027] Step 3:

[2028] Users can access the saved data to replay and analyze past business meetings and interviews.

[2029] Step 4:

[2030] Users can repeatedly practice the areas where they failed or the points they want to improve using simulation videos.

[2031] Specific examples

[2032] For example, if a user conducts a 20-minute online interview, the device captures audio and video data from the microphone and camera in real time and sends it in streaming format to the server. The server analyzes the audio data and calculates the percentage of time the user spoke for 12 minutes and the other party spoke for 8 minutes. This information is overlaid on the user's screen in the form of 68% vs. 32%.

[2033] The server then analyzes the user's facial expressions and tone of voice using an emotion engine, and combines the acquired emotion data to generate an emotion icon. For example, if the server determines that the user is nervous, it will display the emotion icon as "nervous."

[2034] The conversation log is converted into text in real time by a speech recognition engine and displayed on the user's screen. An evaluation score is also calculated in real time, allowing the user to instantly evaluate their own performance.

[2035] When a user uses the quick dictionary feature for a word they don't understand, the meaning of the selected word is displayed in a pop-up, and a suggestion function is also activated to keep the conversation flowing smoothly.

[2036] Finally, after the interview, you can use the retrospective training function to play back past interview data and repeatedly practice areas where you failed or need improvement using simulation footage.

[2037] Example 2

[2038] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2039] In conventional online business meetings and interviews, there was no system that could perform real-time emotion recognition, speech content analysis, evaluation score calculation, next topic suggestion, and even retrospective training all in one place. As a result, users had to rely on multiple tools and manual analysis, making it difficult to improve the quality of efficient communication. In addition, there was little real-time feedback, leading to missed opportunities for improvement.

[2040] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2041] In this invention, the server includes means for acquiring voice data and video data, means for passing the acquired voice data to a voice recognition engine and identifying the content and speaker of the speech, means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions, means for aggregating the speaking time for each speaker and calculating the percentage of speaking time, means for displaying a text-converted conversation log in real time, means for displaying the emotion analysis results as an emotion icon, means for calculating and displaying an evaluation score based on the conversation log and the emotion analysis data, means for analyzing the progress of the conversation and generating and displaying suggestions for what to say next, means for searching for and displaying the meaning of a selected word from a dictionary database, and means for unified management of voice, video, conversation log, and emotion analysis data and for replaying and analyzing past business negotiations and interviews based on the stored data. This allows users to receive real-time and comprehensive data analysis and feedback, significantly improving the quality of online business negotiations and online interviews. 【204...

Claims

1. means for acquiring audio and video data; A means for passing the acquired voice data to a voice recognition engine and identifying the content of the utterance and the speaker; A means for passing the acquired video data to a facial expression analysis engine and identifying emotions from facial expressions; A means for aggregating the speaking time for each speaker and calculating the percentage of time spent speaking; A means to display the textual conversation log in real time, means for displaying the emotion analysis result as an emotion icon; A means for calculating and displaying an evaluation score based on the conversation log and sentiment analysis data; A system including:

2. A means for analyzing the progress of the conversation and generating and displaying suggestions for what to say next; A means to access saved data and replay and analyze past business negotiations and interviews, The system of claim 1 further comprising:

3. 2. The system of claim 1, further comprising means for retrieving and displaying the meaning of a selected word from a dictionary database.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A