System

A system that collects, analyzes, and generates personalized feedback on coaching interactions helps leaders and coaches improve their skills by addressing the lack of objective guidance, enabling continuous skill development.

JP2026018044APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119105
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Leaders and coaches face challenges in providing effective feedback and improving their coaching skills due to a lack of objective and specific guidance in one-on-one and coaching situations, hindering the growth and development of their subordinates.

Method used

A system that collects voice data, converts it into text, analyzes dialogue patterns and emotional tones, generates personalized feedback, and updates its algorithms based on past results to improve analysis accuracy, supporting leaders and coaches in enhancing their skills.

Benefits of technology

Enables leaders and coaches to receive objective and specific feedback, continuously improving their skills and providing effective guidance, leading to enhanced coaching outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018044000001_ABST
    Figure 2026018044000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring voice data; means for converting the acquired voice data into text data; means for analyzing the text data and extracting a conversation pattern and an emotional tone; means for generating feedback based on an analysis result; and means for presenting the generated feedback to a user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Many leaders and coaches face challenges such as not knowing how to have conversations in daily management or one-on-one situations, not developing their subordinates or team members, not improving their coaching skills, and not even knowing if their own methods are appropriate. In such situations, efficient guidance and feedback cannot be provided, hindering the growth of leaders and coaches, so there is a need for a system that identifies areas for improvement in appropriate conversations and provides specific advice tailored to each leader and coach. [Means for solving the problem]

[0005] The present invention is a system that includes a means for acquiring voice data, a means for converting the acquired voice data into text data, a means for analyzing the text data to extract dialogue patterns and emotional tones, a means for generating feedback based on the analysis results, and a means for presenting the generated feedback to a user. The system also includes a means for identifying areas for improvement in open dialogue and active listening based on the analysis results and including such areas in the feedback, and a means for updating the algorithm based on past analysis results and the effects of feedback to improve analysis accuracy, thereby supporting leaders and coaches in effectively improving their skills. This system automatically identifies areas for improvement in conversations and provides individual advice, enabling leaders and coaches to provide effective feedback and promote the improvement of their skills.

[0006] "Audio data" refers to data that records conversations or voices in digital format.

[0007] A "capturing means" is a combination of hardware and software for capturing and storing audio data.

[0008] "Text data" refers to data that is expressed in a transcribed form of audio data.

[0009] "Means for converting" refers to the process and technology for analyzing audio data and converting it into text information.

[0010] "Analysis" and "means of analysis" refer to methods and techniques for extracting meaning and patterns from text data and identifying information.

[0011] A "dialogue pattern" refers to characteristics such as the order and structure of statements made during a conversation, or their frequency.

[0012] "Emotional tone" refers to the specific word choice and vocal inflection that conveys a speaker's feelings or mood in a conversation.

[0013] "Feedback" is information provided based on the analysis results, consisting of areas for improvement, successes, and specific advice.

[0014] "Presentation means" refers to the methods and techniques for visually or audibly displaying feedback to the user.

[0015] "Open dialogue" refers to a form of dialogue in which speakers can freely express their opinions, promoting mutual understanding and empathy.

[0016] "Active listening" is a listening technique that improves the quality of communication through active response to the speaker's words and emotions.

[0017] "Means for updating the algorithm" refers to methods and technologies that evaluate the effectiveness of past analysis results and feedback, and use them to improve the analytical capabilities of the system. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] Embodiments of the present invention will be described in detail below.

[0040] The system of this invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one and coaching sessions. Specifically, it solves the challenges faced by leaders and coaches through a series of processes: collecting, converting, and analyzing audio data, and generating and presenting feedback.

[0041] Program flow:

[0042] Data collection phase:

[0043] 1. Before the start of a 1-on-1 or coaching session, the user sets up the device and launches the recording application, which allows the entire conversation to be recorded.

[0044] 2. The device will start recording as soon as the conversation begins and will capture all audio data during the conversation.

[0045] 3. After the session is over, the user stops it and uploads the recording to the server, which provides the necessary information for the subsequent analysis process.

[0046] Data analysis phase:

[0047] 4. The server converts the received voice data into text data using a voice recognition engine. The voice recognition engine uses an algorithm to convert voice data into text information with high accuracy.

[0048] 5. The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood.

[0049] Feedback generation phase:

[0050] 6. Based on the analysis results, the server identifies areas for improvement and success and generates feedback. For example, if open dialogue is lacking or active listening is not being performed effectively, the feedback will include specific suggestions for improvement.

[0051] 7. After the feedback is generated, the server sends it to the device.

[0052] Results delivery phase:

[0053] 8. The device presents the received feedback to the user, who can then review the feedback and prepare for the next session.

[0054] Examples:

[0055] For example, during an initial one-on-one session, the user starts recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice as feedback. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[0056] This system continuously updates its algorithms based on past analysis results and the effects of feedback, improving analysis accuracy, allowing users to receive more effective feedback with repeated use, enabling leaders and coaches to continually improve their skills.

[0057] The processing flow will be explained below.

[0058] Step 1:

[0059] Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application. The user opens the application and presses the record button to prepare for recording the conversation.

[0060] Step 2:

[0061] The device records audio during the session in real time according to the instructions of the recording application. The device's microphone captures the audio and converts it into a digital audio file. The recorded audio data is temporarily stored on the device's storage.

[0062] Step 3:

[0063] After the session is over, the user stops recording and uploads the recorded data to the server. The user presses the stop recording button and then clicks the "upload data" button, which sends the audio data from the device to the server.

[0064] Step 4:

[0065] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine analyzes the voice data and starts the process of converting it into text information. The converted text data is saved on the server.

[0066] Step 5:

[0067] The server then analyzes the converted text data using natural language processing (NLP) algorithms. During this analysis process, the server identifies dialogue patterns (e.g., order and frequency of statements) and emotional tones (e.g., positive, negative) from the text data. The analysis results are stored on the server for subsequent processing.

[0068] Step 6:

[0069] The server identifies areas for improvement and success based on the analysis results and generates feedback. For example, the server detects a lack of open dialogue or active listening and generates specific improvement advice as feedback. The generated feedback is stored on the server in a structured format (e.g., JSON).

[0070] Step 7:

[0071] The server sends the generated feedback to the device. The feedback data is transferred to the device via API. The transfer method uses encrypted communication for security reasons.

[0072] Step 8:

[0073] The device presents the feedback received from the server to the user, launching a feedback display interface and visually displaying areas for improvement, successes, and specific advice to the user, allowing the user to review this information and prepare for the next session.

[0074] Step 9:

[0075] Based on past analysis results and the effectiveness of feedback, the server updates the algorithm. The server evaluates past data and retrains the machine learning model to improve the system's analysis accuracy. This update occurs periodically, improving analysis accuracy and the quality of feedback.

[0076] Example 1

[0077] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0078] Traditional leadership and coaching feedback tends to be subjective, making it difficult to provide specific and objective feedback. Furthermore, the quality of the feedback is inconsistent, making it difficult to continuously improve skills. This invention solves these problems by providing a system that provides objective and effective feedback based on voice data.

[0079] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0080] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for analyzing the text data and extracting dialogue patterns and emotional tones, means for generating feedback based on the analysis results, and means for providing the generated feedback to the user. This allows leaders and coaches to receive feedback based on specific areas for improvement and successes based on past records, enabling continuous skill development.

[0081] "Audio data" is a recording of the user's speech during a session.

[0082] "Means of collection" refers to the recording application or device used by the user before the session begins, effectively capturing the audio data.

[0083] "Means for converting into text data" refers to a voice recognition engine or software for converting collected voice data into text information.

[0084] "Means for analyzing" refers to the natural language processing (NLP) algorithms and analytical software used to extract dialogue patterns and emotional tone from text data.

[0085] "Means for generating feedback" refers to generative AI models or programs that identify areas for improvement or success based on the analysis results and provide specific advice and evaluation.

[0086] The "means for providing" refers to a communication module or application for notifying the user of the generated feedback and making the content thereof viewable.

[0087] "Open dialogue" is a method of conversation that freely draws out the opinions and thoughts of the other person and encourages active participation.

[0088] "Active listening" is a listening technique in which you listen carefully to what the other person is saying in a conversation and provide appropriate responses and feedback.

[0089] An "algorithm" refers to a series of computational steps or processing flow for solving a specific problem.

[0090] MODE FOR CARRYING OUT THE INVENTION

[0091] An embodiment of the present invention is described in detail below. The system of the present invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one and coaching sessions. Specifically, the system solves the challenges faced by leaders and coaches through a series of processes: collecting, converting, and analyzing audio data, and generating and presenting feedback.

[0092] Hardware and software configuration:

[0093] 1. Data collection phase:

[0094] Before starting a one-on-one or coaching session, the user launches a recording application on a standard device such as a smartphone, tablet, or laptop. Examples of recording applications include Otter.ai and Zoom's recording function.

[0095] By starting the recording application, the terminal enters recording mode and prepares to capture audio data during a conversation.

[0096] 2. Capture audio data:

[0097] The device automatically starts recording as soon as the conversation begins and saves all audio data, which is then continuously recorded until the session ends.

[0098] Users can simply engage in conversation without any special operations, but they can also take notes at important points.

[0099] 3. Uploading recordings:

[0100] After the session is over, the user simply closes the recording application and uploads the recording data to the server via cloud storage (e.g., Google Drive, Dropbox). This operation can also be easily performed through the application interface.

[0101] Data analysis phase:

[0102] 4. Audio to text conversion:

[0103] The server retrieves the voice data from the cloud storage and converts it into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text). This speech recognition engine converts voice into text information using a highly accurate algorithm.

[0104] 5. Text Data Analysis:

[0105] The server analyzes the generated text data using natural language processing (NLP) algorithms (e.g., SpaCy, NLTK). Specifically, it extracts dialogue patterns, analyzes emotional tones, and analyzes the frequency of key phrases and keywords. This analysis makes it possible to understand the flow of the dialogue and the speaker's emotional tendencies.

[0106] Feedback generation phase:

[0107] 6. Feedback Generation:

[0108] Based on the analysis results, the server identifies areas for improvement and success and uses a generative AI model (e.g., OpenAI GPT-3) to generate feedback. This generative model creates specific and personalized feedback based on the analysis results.

[0109] 7. Submitting Feedback:

[0110] Once the feedback is generated, the server sends it to the user's device, usually via email or a dedicated application.

[0111] Results delivery phase:

[0112] 8. Providing feedback:

[0113] The device will notify the user of the received feedback and allow them to review the details of the feedback via a dedicated app or email. Based on the feedback, users can prepare for their next session and implement specific improvements.

[0114] Examples:

[0115] For example, during an initial one-on-one session, the user launches a recording application on their smartphone and begins recording. The device starts recording at the start of the session and records the entire conversation. After the session ends, the user stops recording and uploads the data to online storage using the application's cloud storage function. The server retrieves the audio data from online storage and converts it to text using Google Cloud Speech-to-Text. The converted text data is analyzed using SpaCy to extract dialogue patterns and emotional tones. Based on this, the server generates feedback using the generative AI model GPT-3 and sends the generated feedback to the user's device. Finally, the user receives a notification, can review the feedback in the app, and prepare for the next session.

[0116] Example prompt sentence:

[0117] "I'd like some specific advice on how to improve my active listening in my next one-on-one session."

[0118] "Please give me some examples of specific words I used in past sessions to help improve our Open Dialogue."

[0119] This allows users to receive more detailed and personalized feedback and advice.

[0120] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0121] Step 1:

[0122] Before the start of a 1-on-1 or coaching session, the user launches a recording application. The user launches a dedicated recording application (e.g., general recording software) on a smartphone, tablet, or laptop. This prepares the device to record the entire conversation. The input is the user's operation, and the output is the launch of the recording application.

[0123] Step 2:

[0124] After the recording application is launched, the device enters recording mode and starts capturing audio data as soon as the conversation begins. The entire conversation is recorded and the audio data is saved in storage. The input is the user's speech, and the output is the saved audio data.

[0125] Step 3:

[0126] After the session ends, the user stops the recording application and performs an operation to upload the recorded data to cloud storage (e.g., a general cloud service). The user stops recording and then performs an operation to upload the data to the cloud. The input is stopping the recording application and the audio data, and the output is an audio file saved in cloud storage.

[0127] Step 4:

[0128] The server retrieves the uploaded audio data from the cloud storage. The server automatically accesses the cloud storage and downloads the audio data. The input is the audio file stored in the cloud storage, and the output is the audio file stored on the server.

[0129] Step 5:

[0130] The server converts the voice data into text data using a voice recognition engine (e.g., general voice recognition software). The voice recognition engine analyzes the voice data and converts it into text information. The input is voice data, and the output is the converted text data.

[0131] Step 6:

[0132] The server analyzes the converted text data using natural language processing (NLP) algorithms (e.g., general NLP libraries). Specifically, it extracts dialogue patterns, analyzes emotional tone, extracts key phrases, etc. The input is text data, and the output is the analysis results.

[0133] Step 7:

[0134] The server uses a generative AI model (e.g., a general generative model) to identify areas for improvement and success based on the analysis results and generate feedback. The generative model creates specific and personalized feedback based on the analysis results. The input is the analysis results, and the output is the generated feedback.

[0135] Step 8:

[0136] Once the feedback is generated, the server sends it to the user's device. The generated feedback is sent to the user via email or a dedicated application. The input is the feedback content, and the output is the feedback sent to the user's device.

[0137] Step 9:

[0138] The device notifies the user of the received feedback and allows them to check the details of the feedback via a dedicated application or email. The user prepares for the next session based on the feedback and implements specific improvement measures. The input is the user's feedback confirmation operation, and the output is the user's preparation for the next session.

[0139] (Application example 1)

[0140] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0141] There are limitations to the methods that leaders and coaches have for providing effective feedback to staff in one-on-one or coaching sessions, making it difficult to efficiently and effectively coach and evaluate staff, especially in physical stores. Another problem is that feedback tends to be subjective, making it difficult to accurately grasp the flow of the conversation and the emotional tone. Furthermore, the feedback generated may not directly lead to improvements in staff skills.

[0142] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0143] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for analyzing the text data and extracting dialogue patterns and emotional tones, means for generating feedback based on the analysis results, means for presenting the generated feedback to the user, means for using a natural language processing algorithm to generate the feedback, and means for using these means when training and evaluating staff in physical stores. This makes it possible to provide objective and accurate feedback based on the voice data, which can directly contribute to improving staff skills and training methods.

[0144] "Means for obtaining audio data" refers to the technical methods used to record the voices of leaders and coaches during one-on-one sessions and coaching sessions.

[0145] The "means for converting acquired voice data into text data" refers to a voice recognition technology for converting recorded voice data into text information.

[0146] "Means for analyzing text data and extracting dialogue patterns and emotional tones" refers to a technology that uses natural language processing algorithms to analyze conversation content and identify the content of statements and emotional trends.

[0147] The "means for generating feedback based on the analysis results" is a method for generating feedback based on the analyzed data that specifically indicates how the user can improve their skills and areas for improvement.

[0148] The "means for presenting the generated feedback to the user" is a technology for providing the generated feedback to the leader or coach visually or audibly.

[0149] A "natural language processing algorithm" is a computational method for understanding and analyzing text data and extracting information based on its content.

[0150] "Means of using these methods when instructing and evaluating staff in physical stores" refers to methods for leaders and coaches to use the above-mentioned technological methods when instructing and evaluating staff in physical stores.

[0151] The system of the present invention is designed to enable leaders and coaches to provide effective feedback to staff in one-on-one sessions and coaching sessions in brick-and-mortar stores. Specifically, it collects voice data, converts it into text data, and then analyzes it to generate feedback, which is then presented to the user.

[0152] Hardware and software used

[0153] The main components of this system include a smartphone, a server, a natural language processing (NLP) engine, and a voice recognition engine.

[0154] Smartphone: A device used by leaders and coaches to record and transmit audio data.

[0155] Server: A central computer system that receives and analyzes voice data.

[0156] Speech recognition engine: Software for converting voice data into text data.

[0157] Natural Language Processing Engine (NLP Engine): Provides algorithms for analyzing text data and extracting dialogue patterns and emotional tone.

[0158] Data processing and calculation

[0159] 1. The smartphone records the audio as soon as the conversation begins and uploads the conversation data to a server.

[0160] 2. The server converts the received voice data into text data using a speech recognition engine. This conversion uses a highly accurate algorithm.

[0161] 3. The converted text data is analyzed by an NLP engine on the server, which extracts dialogue patterns and emotional tones.

[0162] 4. Based on the extracted data, the server generates feedback, for example, suggesting improvements to the open dialogue if the dialogue tends to be one-way.

[0163] 5. The generated feedback is then sent back to the smartphone and presented to the leader or coach.

[0164] Specific examples

[0165] For example, consider a one-on-one session between a new staff member and the store manager at a physical store. The store manager launches an application and records the conversation. After the recording is complete, the data is uploaded to a server. The server converts the audio data into text data and analyzes it. Based on the analysis results, feedback such as "You should practice more active listening" is generated and displayed on the smartphone.

[0166] Prompt Sentence Examples

[0167] Here are some examples of prompts the system might use:

[0168] "Please analyze the following conversation logs to extract emotional tone and areas for improvement. Additionally, please generate and provide specific feedback."

[0169] Interaction logs:

[0170] Manager: "How was work today?"

[0171] Staff: "It was hard work, but it was fun."

[0172] Feedback generated:

[0173] "The positive response was very strong. Please continue to support our staff in this way next time."

[0174] This system allows leaders and coaches to receive objective and accurate feedback, which directly contributes to improving staff skills and teaching methods.

[0175] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0176] Step 1:

[0177] Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application, which prepares the device to record the conversation. The input is the start event of the session, and the output is the status that recording is ready.

[0178] Step 2:

[0179] The device starts recording as soon as the conversation begins and captures all audio data during the conversation. The input is the audio data from the session, and the output is the recorded audio file, which is used for the subsequent analysis process.

[0180] Step 3:

[0181] After the session is over, the user stops recording and uploads the recording to the server. The input is the recorded audio file, and the output is the audio data stored on the server.

[0182] Step 4:

[0183] The server converts the received voice data into text data using a voice recognition engine. The voice recognition engine uses an algorithm to convert voice data into text information with high accuracy. The input is the uploaded voice data, and the output is the converted text data.

[0184] Step 5:

[0185] The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood. The input is text data, and the output is the analysis results of dialogue patterns and emotional tones.

[0186] Step 6:

[0187] Based on the analysis results, the server identifies areas for improvement and success and generates feedback. For example, if open dialogue is lacking or active listening is not being performed effectively, the feedback will include specific suggestions for improvement. The input is the analysis results of dialogue patterns and emotional tone, and the output is the generated feedback.

[0188] Step 7:

[0189] The generated feedback is sent from the server to the terminal, where the input is the generated feedback and the output is the feedback information presented to the user.

[0190] Step 8:

[0191] The terminal presents the received feedback to the user, who then checks the feedback and prepares for the next session. The input is the feedback information sent from the server, and the output is a feedback display that the user can check.

[0192] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0193] Embodiments of the present invention will be described in detail below.

[0194] The system of the present invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one and coaching sessions. Specifically, it solves the challenges faced by leaders and coaches through a series of processes: collecting, converting, analyzing, and recognizing emotions in audio data, as well as generating and presenting feedback.

[0195] Program flow:

[0196] Data collection phase:

[0197] 1. Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application. The user opens the application and presses the record button to prepare for recording the conversation.

[0198] 2. The device will record audio during the session in real time as instructed by the recording application. The device's microphone will capture the audio and convert it into a digital audio file. The recorded audio data will be temporarily stored on the device's storage.

[0199] 3. After the session is over, the user stops the recording and uploads the recording data to the server. The user presses the stop recording button and then clicks the "Upload Data" button, which sends the audio data from the device to the server.

[0200] Data analysis phase:

[0201] 4. The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine analyzes the voice data and starts the process of converting it into text information. The converted text data is stored on the server.

[0202] 5. The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood.

[0203] Emotion Recognition Phase:

[0204] 6. The analyzed text and voice data are then subjected to emotion analysis by the emotion engine on the server. The emotion engine recognizes emotions from the tone of voice and text and identifies the user's emotional state. This emotion data is used to generate feedback.

[0205] Feedback generation phase:

[0206] 7. Based on the analysis results and emotion recognition results, the server identifies areas for improvement and success and generates feedback. For example, the server analyzes the lack of open dialogue, lack of active listening, and appropriate emotional tone, and generates specific improvement advice as feedback. The generated feedback is stored on the server in a structured format (e.g., JSON).

[0207] 8. After the feedback is generated, the server sends it to the device. The feedback data is transferred to the device via API. The transfer method is encrypted for security reasons.

[0208] Results delivery phase:

[0209] 9. The device presents the feedback received from the server to the user. It launches a feedback display interface and visually displays areas for improvement, successes, and specific advice to the user. The user can review this information and prepare for the next session.

[0210] Examples:

[0211] For example, during an initial one-on-one session, the user begins recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results and emotion recognition by the emotion engine, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice as feedback. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[0212] This system continuously updates its algorithms based on past analysis results and the effects of feedback, improving analysis accuracy, allowing users to receive more effective feedback with repeated use, enabling leaders and coaches to continually improve their skills.

[0213] The processing flow will be explained below.

[0214] Step 1:

[0215] Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application. The user opens the application and is ready to press the record button to start the recording function.

[0216] Step 2:

[0217] The device records audio during the session in real time as directed by the recording application. The device's microphone captures the audio and stores it as a digital audio file. The recorded audio data is temporarily stored on the device's storage.

[0218] Step 3:

[0219] After the session is over, the user stops recording and uploads the recorded data to the server. The user presses the stop recording button and then clicks the "upload data" button, which sends the audio data from the device to the server.

[0220] Step 4:

[0221] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine analyzes the voice data and converts it into text information. The converted text data is stored on the server.

[0222] Step 5:

[0223] The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood. The analysis results are stored on the server and used for subsequent processing.

[0224] Step 6:

[0225] The analyzed text and voice data are then subjected to emotion analysis by an emotion engine within the server. The emotion engine recognizes emotions from voice tone and text and identifies the user's emotional state. This emotion data is then used to generate feedback.

[0226] Step 7:

[0227] Based on the analysis results and emotion recognition results, the server identifies areas for improvement and success and generates feedback. For example, the server analyzes the lack of open dialogue, lack of active listening, and appropriate emotional tone, and generates specific improvement advice as feedback. The generated feedback is stored on the server in a structured format (e.g., JSON).

[0228] Step 8:

[0229] After the feedback is generated, the server sends it to the device. The feedback data is transferred to the device via API. The transfer method uses encrypted communication to ensure security.

[0230] Step 9:

[0231] The device presents the feedback received from the server to the user, launching a feedback display interface and visually displaying areas for improvement, successes, and specific advice to the user, allowing the user to review this information and prepare for the next session.

[0232] Step 10:

[0233] Based on past analysis results and the effectiveness of feedback, the server updates the algorithm. The server evaluates past data and retrains the machine learning model to improve the system's analysis accuracy. This update occurs periodically, improving analysis accuracy and the quality of feedback.

[0234] Examples:

[0235] For example, during an initial one-on-one session, the user begins recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results and emotion recognition by the emotion engine, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice as feedback. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[0236] This system continuously updates its algorithms based on past analysis results and the effects of feedback, improving analysis accuracy, allowing users to receive more effective feedback with repeated use, enabling leaders and coaches to continually improve their skills.

[0237] Example 2

[0238] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0239] For leaders and coaches to provide effective feedback during one-on-one or coaching sessions, they need to accurately record the conversations that occur during the session and analyze the dialogue patterns and emotional tone. However, manually analyzing this information takes time and effort, and the accuracy of the analysis is often insufficient. Furthermore, if the quality of feedback does not improve, there is a problem that the skill development of leaders and coaches stagnates. The present invention aims to solve these problems and provide a system for improving the effectiveness of sessions.

[0240] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0241] In this invention, the server includes: means for a user to set up a terminal and launch a recording application before a one-on-one session or a coaching session; means for the terminal to record audio during the session in real time according to instructions from the recording application; means for the user to upload the recorded data to the server after the session ends; means for the server to pass the received audio data to a speech recognition engine and convert it into text data; means for the server to analyze the converted text data and extract dialogue patterns and emotional tones; means for an emotion engine in the server to perform emotion analysis on the analyzed text data and audio data; means for the server to generate feedback based on the analysis results and emotion recognition results; means for using a generation AI model when generating feedback; and means for presenting the generated feedback to the user. This automates the process from recording a session to analyzing it and providing feedback, making it possible to provide effective feedback quickly and accurately.

[0242] "User" refers to the person or parties who operate the recording application to collect data and review feedback when conducting one-on-one sessions or coaching sessions.

[0243] "Device" refers to the electronic device on which the user launches the recording application, collects audio data, and displays feedback, such as a smartphone or PC.

[0244] "Recording application" refers to software that uses a device to record audio from one-on-one sessions or coaching sessions in real time and save it as a digital audio file.

[0245] "Server" refers to a computer system that receives recorded voice data, analyzes the data using a voice recognition engine and emotion engine, and generates feedback.

[0246] "Speech recognition engine" refers to a technology or software module that analyzes voice data received by a server and converts it into text data.

[0247] "Text data" refers to character information converted from voice data by a voice recognition engine.

[0248] "Dialogue patterns" refer to the structure and flow of conversation extracted by analyzing text data.

[0249] "Emotional tone" refers to the emotional tendencies and tone of a speaker in a conversation.

[0250] "Emotion engine" refers to a technology or software module that performs emotion analysis based on analyzed text data and voice data to identify the user's emotional state.

[0251] "Feedback" refers to advice and specific guidelines for improvement generated based on analysis results and emotion recognition results.

[0252] A "generative AI model" is an artificial intelligence model used by the server to generate feedback, and refers to a technology that uses generative AI to generate specific improvement advice, for example.

[0253] The system of the present invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one sessions and coaching sessions. Specifically, it includes the following operations:

[0254] Data Collection Phase

[0255] First, users set up their device before a one-on-one or coaching session and launch the recording application. Users open the application on their smartphone, tablet, or computer and press the record button to begin recording the session.

[0256] The device, in accordance with the instructions of the recording application, records audio during the session in real time using the built-in microphone and converts it into a digital audio file, which is then temporarily stored in the device's storage device.

[0257] After the session ends, the user stops recording and uploads the recorded data to the server. Specifically, when the user clicks the "Stop Recording" button and then the "Upload Data" button, the audio data is sent from the device to the server using encrypted communication.

[0258] Data analysis phase

[0259] The server passes the received voice data to a speech recognition engine and converts it into text data. This speech recognition uses a common speech recognition service (e.g., Google Cloud Speech-to-Text API). The converted text data is stored in a database on the server.

[0260] The server then analyzes the converted text data and applies natural language processing (NLP) algorithms (e.g., SpaCy or NLTK) to extract dialogue patterns and emotional tones, thereby understanding the flow of dialogue and the speaker's emotional tendencies.

[0261] Emotion Recognition Phase

[0262] Based on the analyzed text and voice data, the emotion engine in the server performs emotion analysis. The emotion engine uses a deep learning model to identify the user's emotional state from the tone of voice and text. This allows it to classify emotions as positive, negative, or neutral.

[0263] Feedback generation phase

[0264] The server generates feedback based on the analysis results and emotion recognition results. This feedback is generated using a "generative AI model" (e.g., OpenAI GPT-3) to generate specific advice based on the analysis data. The generated feedback is stored on the server in a structured format (e.g., JSON).

[0265] After the feedback is generated, the server sends it to the device. The feedback data is transmitted via an API using a secure communication protocol.

[0266] Results delivery phase

[0267] The device then presents the feedback received from the server to the user, who is then presented with a dedicated application or web interface that visually displays areas for improvement, successes, and specific advice.The user can then review this information and prepare for the next session.

[0268] Specific examples

[0269] For example, during an initial one-on-one session, the user starts recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results and emotion recognition by the emotion engine, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[0270] Prompt Sentence Examples

[0271] "Please analyze the following audio data and generate feedback on problems and areas for improvement:"

[0272] "All recordings of the first 1-on-1 session"

[0273] "Analyze dialogue patterns and emotional tone"

[0274] "Generate feedback with specific advice on improving lack of open dialogue and active listening"

[0275] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0276] Step 1:

[0277] The user sets up the device and launches the recording application before the one-on-one or coaching session, which involves the user opening the dedicated recording application on their smartphone, tablet, or PC and clicking the start button.

[0278] Input: A recording application that is initiated by user action.

[0279] Output: Ready for recording.

[0280] Step 2:

[0281] The device will record audio during the session in real time as directed by the recording application, using the device's built-in microphone to convert the audio into a digital audio file that is immediately saved.

[0282] Input: Audio during the session.

[0283] Output: A digital audio file is temporarily saved to the device storage.

[0284] Step 3:

[0285] After the session ends, the user stops recording and uploads the recorded data to the server. Specifically, the user presses the application's stop recording button and then clicks the "upload data" button, which sends the audio data from the device to the server. The data is transferred using encrypted communication.

[0286] Input: Recorded audio file.

[0287] Output: The audio data is uploaded to the server.

[0288] Step 4:

[0289] The server passes the received voice data to a voice recognition engine and converts it into text data. This process includes sending the voice data to a voice recognition service (e.g., a voice recognition API) and outputting it as text data.

[0290] Input: Recorded audio data.

[0291] Output: Text data is generated and stored on the server.

[0292] Step 5:

[0293] The server analyzes the converted text data to extract dialogue patterns and emotional tones. Natural language processing (NLP) algorithms (e.g., SpaCy and NLTK) are applied to analyze the text data to identify conversation flow and emotional trends.

[0294] Input: Text data converted by speech recognition.

[0295] Output: Information on dialogue patterns and emotional tone.

[0296] Step 6:

[0297] Based on the analyzed text and voice data, the emotion engine in the server performs emotion analysis. The emotion engine uses a deep learning model to analyze the user's emotional state from voice tone and text.

[0298] Input: Information on interaction patterns and emotional tone.

[0299] Output: Sentiment classification data such as positive, negative, or neutral.

[0300] Step 7:

[0301] The server generates specific feedback based on the analysis results and emotion recognition results. This feedback is generated using a generative AI model (e.g., a generative AI model). The server creates specific improvement advice based on the analysis data and saves it in JSON format.

[0302] Input: Sentiment classification data and dialogue pattern analysis data.

[0303] Output: Feedback data is generated.

[0304] Step 8:

[0305] The server sends the generated feedback data to the device via an API using an encrypted communication protocol.

[0306] Input: Feedback data.

[0307] Output: Feedback data is sent to the terminal.

[0308] Step 9:

[0309] The device receives feedback from the server and presents it to the user. Using a dedicated application or web interface, the user is visually shown areas for improvement, successes, and specific advice. This allows the user to prepare for the next session.

[0310] Input: Feedback data.

[0311] Output: Feedback is presented to the user.

[0312] (Application example 2)

[0313] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0314] Currently, it is difficult to systematically and efficiently provide training and feedback to improve customer service skills in brick-and-mortar stores. To improve employees' customer service skills, a system is needed to evaluate their interactions with customers during their daily work and provide specific feedback on areas for improvement. Such a system must also precisely analyze emotional tone and conversation patterns to evaluate positive tone and customer service.

[0315] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0316] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for analyzing the text data and extracting dialogue patterns and emotional tones, means for evaluating customer service skills based on the analysis results and emotional tones, means for providing specific points for improvement in customer service as feedback, means for presenting the generated feedback to the user, means for updating the algorithm based on past analysis results and the effects of the feedback to improve analysis accuracy, and means for using a generative AI model to provide feedback and input prompt sentences. This makes it possible to efficiently and accurately provide specific points for improvement to improve the customer service skills of employees.

[0317] "Voice data" is information that is a digital recording of what a user says.

[0318] "Text data" is character string information obtained by analyzing voice data.

[0319] A "dialogue pattern" is information that indicates the flow and structure of a conversation, including who spoke, when, and the content of the conversation.

[0320] "Emotional tone" is information obtained by analyzing the speaker's emotional state and emotional expression during a conversation.

[0321] "Feedback" refers to specific improvements and advice regarding a user's behavior and performance that is generated based on analysis results and evaluations.

[0322] "Customer service skills" refers to the abilities and techniques that employees have when interacting with and providing service to customers in physical stores.

[0323] A "positive tone" refers to a positive and friendly manner of speaking and behavior.

[0324] "Positive customer service" refers to the act and technique of treating customers in a friendly and cheerful manner.

[0325] "Analysis precision" refers to the degree of accuracy and reproducibility of data analysis.

[0326] A "generative AI model" is a model that uses artificial intelligence technology to generate a model tailored to a specific task, and is used for data analysis and feedback generation.

[0327] A "prompt sentence" is a sentence used to input specific instructions or questions to a generative AI model.

[0328] This invention is a system for improving customer service skills in brick-and-mortar stores, and provides a series of processes for acquiring, analyzing, and generating feedback using voice data. This system has the following main hardware and software configuration:

[0329] System hardware and software configuration

[0330] 1. User device (smartphone)

[0331] Microphone: Captures audio data

[0332] Recording application: An application for recording audio data and uploading it to a server.

[0333] 2. Server

[0334] Speech recognition engine: converts voice data into text data (e.g., Google Speech Recognition)

[0335] Natural Language Processing (NLP) algorithms: Analyze text data and extract dialogue patterns and emotional tone (e.g., TextBlob)

[0336] Emotion Recognition Engine: Recognizes emotions from voice tone and text (e.g. TextBlob)

[0337] Generative AI model: Generates feedback based on analysis results

[0338] Database: Stores text data and feedback results

[0339] 3. Feedback display interface

[0340] User Interface: Present the generated feedback to the user in a GUI

[0341] Data processing and calculation

[0342] 1. Acquiring audio data

[0343] The user's smartphone records conversations with customers through a microphone.

[0344] The recording application converts the audio data into a digital format in real time and temporarily stores it in the device's storage.

[0345] 2. Uploading audio data

[0346] After finishing recording, the user clicks the "Upload Data" button and the audio data is sent to the server.

[0347] 3. Speech Recognition and Text Conversion

[0348] The server passes the received voice data to a voice recognition engine, which converts the voice data into text data.

[0349] 4. Text Data Analysis

[0350] The converted text data is then used to extract dialogue patterns and emotional tones using natural language processing (NLP) algorithms, which allow for understanding the flow of the dialogue and the speaker's emotional tendencies.

[0351] 5. Emotion analysis

[0352] The analyzed text and voice data are then subjected to emotion analysis by an emotion recognition engine, which recognizes emotions from voice tone and text to identify the user's emotional state.

[0353] 6. Feedback Generation

[0354] Based on the analysis and emotion recognition results, the server generates feedback on customer service skills and specific areas for improvement. For example, the server can determine whether there is a lack of open dialogue, a lack of active listening, or appropriate emotional tone, and provide specific advice.

[0355] 7. Providing Feedback

[0356] The generated feedback is sent from the server to the user's device and presented to the user through the device's feedback display interface, allowing the user to improve their skills based on this feedback.

[0357] Examples of concrete examples and prompts

[0358] 1. Example:

[0359] Employee A starts recording with the app while talking to Customer B.

[0360] After the conversation is over, press the "Analyze" button in the app to upload the recording data to the server.

[0361] The server converts the speech into text and performs sentiment analysis.

[0362] Based on the results of the sentiment analysis, feedback such as "Try speaking in a more positive tone" is generated.

[0363] Employee A will take that feedback into consideration at the start of their next shift.

[0364] 2. Example prompt for the generative AI model:

[0365] Prompt: Convert the following audio data into text and perform sentiment analysis. Generate appropriate feedback. Audio data: <Audio data link>

[0366] In this way, the inventive system provides specific improvements for efficiently and effectively improving employee customer service skills.

[0367] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0368] Step 1:

[0369] The user launches the recording application and presses the start recording button. The input is the user's operation, and the output is the transition to recording mode. The recording application captures the voice data during the customer interaction with the device's microphone and temporarily stores it as a digital audio file in the device's storage.

[0370] Step 2:

[0371] After finishing recording, the user presses the stop recording button. The input is the user's operation, and the output is stopping the recording. Next, the user presses the "upload data" button to send the audio data to the server. The input is the audio file, and the output is the audio file uploaded to the server.

[0372] Step 3:

[0373] The server passes the received voice data to a speech recognition engine and converts the voice file into text data. The input is the voice file, and the output is the converted text data. This process uses the Google Speech Recognition API or similar to convert voice to text.

[0374] Step 4:

[0375] The server passes the text data to a natural language processing (NLP) algorithm to extract dialogue patterns and emotional tones. The input is the text data, and the output is the analyzed dialogue patterns and emotional tones. This analysis is performed using tools such as TextBlob.

[0376] Step 5:

[0377] The server passes the analyzed text and voice data to the emotion recognition engine for emotion analysis. The input is the analyzed text and voice data, and the output is the user's emotional state data. The emotion engine analyzes the emotional tone and identifies emotions such as positive, negative, and neutral.

[0378] Step 6:

[0379] The server generates feedback using a generative AI model based on the analysis results and emotion recognition results. The inputs are dialogue patterns, emotional tone, and emotional state, and the output is specific feedback. For example, the generative AI model generates feedback such as "Try speaking in a more positive tone." The generative AI model is given instructions using prompt sentences.

[0380] Step 7:

[0381] The server sends the generated feedback to the terminal. The input is the generated feedback, and the output is the transmission of the feedback data to the terminal. The feedback data is securely transmitted using encrypted communication.

[0382] Step 8:

[0383] The terminal presents the feedback received from the server to the user. The input is feedback data, and the output is feedback displayed to the user. The feedback display interface visually presents the feedback content to the user, who can review it and use it in the next session to improve their customer service skills.

[0384] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0385] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0386] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0387] [Second embodiment]

[0388] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0389] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0390] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0391] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0392] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0393] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0394] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0395] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0396] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0397] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0398] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0399] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0400] Embodiments of the present invention will be described in detail below.

[0401] The system of this invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one and coaching sessions. Specifically, it solves the challenges faced by leaders and coaches through a series of processes: collecting, converting, and analyzing audio data, and generating and presenting feedback.

[0402] Program flow:

[0403] Data collection phase:

[0404] 1. Before the start of a 1-on-1 or coaching session, the user sets up the device and launches the recording application, which allows the entire conversation to be recorded.

[0405] 2. The device will start recording as soon as the conversation begins and will capture all audio data during the conversation.

[0406] 3. After the session is over, the user stops it and uploads the recording to the server, which provides the necessary information for the subsequent analysis process.

[0407] Data analysis phase:

[0408] 4. The server converts the received voice data into text data using a voice recognition engine. The voice recognition engine uses an algorithm to convert voice data into text information with high accuracy.

[0409] 5. The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood.

[0410] Feedback generation phase:

[0411] 6. Based on the analysis results, the server identifies areas for improvement and success and generates feedback. For example, if open dialogue is lacking or active listening is not being performed effectively, the feedback will include specific suggestions for improvement.

[0412] 7. After the feedback is generated, the server sends it to the device.

[0413] Results delivery phase:

[0414] 8. The device presents the received feedback to the user, who can then review the feedback and prepare for the next session.

[0415] Examples:

[0416] For example, during an initial one-on-one session, the user starts recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice as feedback. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[0417] This system continuously updates its algorithms based on past analysis results and the effects of feedback, improving analysis accuracy, allowing users to receive more effective feedback with repeated use, enabling leaders and coaches to continually improve their skills.

[0418] The processing flow will be explained below.

[0419] Step 1:

[0420] Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application. The user opens the application and presses the record button to prepare for recording the conversation.

[0421] Step 2:

[0422] The device records audio during the session in real time according to the instructions of the recording application. The device's microphone captures the audio and converts it into a digital audio file. The recorded audio data is temporarily stored on the device's storage.

[0423] Step 3:

[0424] After the session is over, the user stops recording and uploads the recorded data to the server. The user presses the stop recording button and then clicks the "upload data" button, which sends the audio data from the device to the server.

[0425] Step 4:

[0426] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine analyzes the voice data and starts the process of converting it into text information. The converted text data is saved on the server.

[0427] Step 5:

[0428] The server then analyzes the converted text data using natural language processing (NLP) algorithms. During this analysis process, the server identifies dialogue patterns (e.g., order and frequency of statements) and emotional tones (e.g., positive, negative) from the text data. The analysis results are stored on the server for subsequent processing.

[0429] Step 6:

[0430] The server identifies areas for improvement and success based on the analysis results and generates feedback. For example, the server detects a lack of open dialogue or active listening and generates specific improvement advice as feedback. The generated feedback is stored on the server in a structured format (e.g., JSON).

[0431] Step 7:

[0432] The server sends the generated feedback to the device. The feedback data is transferred to the device via API. The transfer method uses encrypted communication for security reasons.

[0433] Step 8:

[0434] The device presents the feedback received from the server to the user, launching a feedback display interface and visually displaying areas for improvement, successes, and specific advice to the user, allowing the user to review this information and prepare for the next session.

[0435] Step 9:

[0436] Based on past analysis results and the effectiveness of feedback, the server updates the algorithm. The server evaluates past data and retrains the machine learning model to improve the system's analysis accuracy. This update occurs periodically, improving analysis accuracy and the quality of feedback.

[0437] Example 1

[0438] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0439] Traditional leadership and coaching feedback tends to be subjective, making it difficult to provide specific and objective feedback. Furthermore, the quality of the feedback is inconsistent, making it difficult to continuously improve skills. This invention solves these problems by providing a system that provides objective and effective feedback based on voice data.

[0440] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0441] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for analyzing the text data and extracting dialogue patterns and emotional tones, means for generating feedback based on the analysis results, and means for providing the generated feedback to the user. This allows leaders and coaches to receive feedback based on specific areas for improvement and successes based on past records, enabling continuous skill development.

[0442] "Audio data" is a recording of the user's speech during a session.

[0443] "Means of collection" refers to the recording application or device used by the user before the session begins, effectively capturing the audio data.

[0444] "Means for converting into text data" refers to a voice recognition engine or software for converting collected voice data into text information.

[0445] "Means for analyzing" refers to the natural language processing (NLP) algorithms and analytical software used to extract dialogue patterns and emotional tone from text data.

[0446] "Means for generating feedback" refers to generative AI models or programs that identify areas for improvement or success based on the analysis results and provide specific advice and evaluation.

[0447] The "means for providing" refers to a communication module or application for notifying the user of the generated feedback and making the content thereof viewable.

[0448] "Open dialogue" is a method of conversation that freely draws out the opinions and thoughts of the other person and encourages active participation.

[0449] "Active listening" is a listening technique in which you listen carefully to what the other person is saying in a conversation and provide appropriate responses and feedback.

[0450] An "algorithm" refers to a series of computational steps or processing flow for solving a specific problem.

[0451] MODE FOR CARRYING OUT THE INVENTION

[0452] An embodiment of the present invention is described in detail below. The system of the present invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one and coaching sessions. Specifically, the system solves the challenges faced by leaders and coaches through a series of processes: collecting, converting, and analyzing audio data, and generating and presenting feedback.

[0453] Hardware and software configuration:

[0454] 1. Data collection phase:

[0455] Before starting a one-on-one or coaching session, the user launches a recording application on a standard device such as a smartphone, tablet, or laptop. Examples of recording applications include Otter.ai and Zoom's recording function.

[0456] By starting the recording application, the terminal enters recording mode and prepares to capture audio data during a conversation.

[0457] 2. Capture audio data:

[0458] The device automatically starts recording as soon as the conversation begins and saves all audio data, which is then continuously recorded until the session ends.

[0459] Users can simply engage in conversation without any special operations, but they can also take notes at important points.

[0460] 3. Uploading recordings:

[0461] After the session is over, the user simply closes the recording application and uploads the recording data to the server via cloud storage (e.g., Google Drive, Dropbox). This operation can also be easily performed through the application interface.

[0462] Data analysis phase:

[0463] 4. Audio to text conversion:

[0464] The server retrieves the voice data from the cloud storage and converts it into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text). This speech recognition engine converts voice into text information using a highly accurate algorithm.

[0465] 5. Text Data Analysis:

[0466] The server analyzes the generated text data using natural language processing (NLP) algorithms (e.g., SpaCy, NLTK). Specifically, it extracts dialogue patterns, analyzes emotional tones, and analyzes the frequency of key phrases and keywords. This analysis makes it possible to understand the flow of the dialogue and the speaker's emotional tendencies.

[0467] Feedback generation phase:

[0468] 6. Feedback Generation:

[0469] Based on the analysis results, the server identifies areas for improvement and success and uses a generative AI model (e.g., OpenAI GPT-3) to generate feedback. This generative model creates specific and personalized feedback based on the analysis results.

[0470] 7. Submitting Feedback:

[0471] Once the feedback is generated, the server sends it to the user's device, usually via email or a dedicated application.

[0472] Results delivery phase:

[0473] 8. Providing feedback:

[0474] The device will notify the user of the received feedback and allow them to review the details of the feedback via a dedicated app or email. Based on the feedback, users can prepare for their next session and implement specific improvements.

[0475] Examples:

[0476] For example, during an initial one-on-one session, the user launches a recording application on their smartphone and begins recording. The device starts recording at the start of the session and records the entire conversation. After the session ends, the user stops recording and uploads the data to online storage using the application's cloud storage function. The server retrieves the audio data from online storage and converts it to text using Google Cloud Speech-to-Text. The converted text data is analyzed using SpaCy to extract dialogue patterns and emotional tones. Based on this, the server generates feedback using the generative AI model GPT-3 and sends the generated feedback to the user's device. Finally, the user receives a notification, can review the feedback in the app, and prepare for the next session.

[0477] Example prompt sentence:

[0478] "I'd like some specific advice on how to improve my active listening in my next one-on-one session."

[0479] "Please give me some examples of specific words I used in past sessions to help improve our Open Dialogue."

[0480] This allows users to receive more detailed and personalized feedback and advice.

[0481] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0482] Step 1:

[0483] Before the start of a 1-on-1 or coaching session, the user launches a recording application. The user launches a dedicated recording application (e.g., general recording software) on a smartphone, tablet, or laptop. This prepares the device to record the entire conversation. The input is the user's operation, and the output is the launch of the recording application.

[0484] Step 2:

[0485] After the recording application is launched, the device enters recording mode and starts capturing audio data as soon as the conversation begins. The entire conversation is recorded and the audio data is saved in storage. The input is the user's speech, and the output is the saved audio data.

[0486] Step 3:

[0487] After the session ends, the user stops the recording application and performs an operation to upload the recorded data to cloud storage (e.g., a general cloud service). The user stops recording and then performs an operation to upload the data to the cloud. The input is stopping the recording application and the audio data, and the output is an audio file saved in cloud storage.

[0488] Step 4:

[0489] The server retrieves the uploaded audio data from the cloud storage. The server automatically accesses the cloud storage and downloads the audio data. The input is the audio file stored in the cloud storage, and the output is the audio file stored on the server.

[0490] Step 5:

[0491] The server converts the voice data into text data using a voice recognition engine (e.g., general voice recognition software). The voice recognition engine analyzes the voice data and converts it into text information. The input is voice data, and the output is the converted text data.

[0492] Step 6:

[0493] The server analyzes the converted text data using natural language processing (NLP) algorithms (e.g., general NLP libraries). Specifically, it extracts dialogue patterns, analyzes emotional tone, extracts key phrases, etc. The input is text data, and the output is the analysis results.

[0494] Step 7:

[0495] The server uses a generative AI model (e.g., a general generative model) to identify areas for improvement and success based on the analysis results and generate feedback. The generative model creates specific and personalized feedback based on the analysis results. The input is the analysis results, and the output is the generated feedback.

[0496] Step 8:

[0497] Once the feedback is generated, the server sends it to the user's device. The generated feedback is sent to the user via email or a dedicated application. The input is the feedback content, and the output is the feedback sent to the user's device.

[0498] Step 9:

[0499] The device notifies the user of the received feedback and allows them to check the details of the feedback via a dedicated application or email. The user prepares for the next session based on the feedback and implements specific improvement measures. The input is the user's feedback confirmation operation, and the output is the user's preparation for the next session.

[0500] (Application example 1)

[0501] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0502] There are limitations to the methods that leaders and coaches have for providing effective feedback to staff in one-on-one or coaching sessions, making it difficult to efficiently and effectively coach and evaluate staff, especially in physical stores. Another problem is that feedback tends to be subjective, making it difficult to accurately grasp the flow of the conversation and the emotional tone. Furthermore, the feedback generated may not directly lead to improvements in staff skills.

[0503] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0504] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for analyzing the text data and extracting dialogue patterns and emotional tones, means for generating feedback based on the analysis results, means for presenting the generated feedback to the user, means for using a natural language processing algorithm to generate the feedback, and means for using these means when training and evaluating staff in physical stores. This makes it possible to provide objective and accurate feedback based on the voice data, which can directly contribute to improving staff skills and training methods.

[0505] "Means for obtaining audio data" refers to the technical methods used to record the voices of leaders and coaches during one-on-one sessions and coaching sessions.

[0506] The "means for converting acquired voice data into text data" refers to a voice recognition technology for converting recorded voice data into text information.

[0507] "Means for analyzing text data and extracting dialogue patterns and emotional tones" refers to a technology that uses natural language processing algorithms to analyze conversation content and identify the content of statements and emotional trends.

[0508] The "means for generating feedback based on the analysis results" is a method for generating feedback based on the analyzed data that specifically indicates how the user can improve their skills and areas for improvement.

[0509] The "means for presenting the generated feedback to the user" is a technology for providing the generated feedback to the leader or coach visually or audibly.

[0510] A "natural language processing algorithm" is a computational method for understanding and analyzing text data and extracting information based on its content.

[0511] "Means of using these methods when instructing and evaluating staff in physical stores" refers to methods for leaders and coaches to use the above-mentioned technological methods when instructing and evaluating staff in physical stores.

[0512] The system of the present invention is designed to enable leaders and coaches to provide effective feedback to staff in one-on-one sessions and coaching sessions in brick-and-mortar stores. Specifically, it collects voice data, converts it into text data, and then analyzes it to generate feedback, which is then presented to the user.

[0513] Hardware and software used

[0514] The main components of this system include a smartphone, a server, a natural language processing (NLP) engine, and a voice recognition engine.

[0515] Smartphone: A device used by leaders and coaches to record and transmit audio data.

[0516] Server: A central computer system that receives and analyzes voice data.

[0517] Speech recognition engine: Software for converting voice data into text data.

[0518] Natural Language Processing Engine (NLP Engine): Provides algorithms for analyzing text data and extracting dialogue patterns and emotional tone.

[0519] Data processing and calculation

[0520] 1. The smartphone records the audio as soon as the conversation begins and uploads the conversation data to a server.

[0521] 2. The server converts the received voice data into text data using a speech recognition engine. This conversion uses a highly accurate algorithm.

[0522] 3. The converted text data is analyzed by an NLP engine on the server, which extracts dialogue patterns and emotional tones.

[0523] 4. Based on the extracted data, the server generates feedback, for example, suggesting improvements to the open dialogue if the dialogue tends to be one-way.

[0524] 5. The generated feedback is then sent back to the smartphone and presented to the leader or coach.

[0525] Specific examples

[0526] For example, consider a one-on-one session between a new staff member and the store manager at a physical store. The store manager launches an application and records the conversation. After the recording is complete, the data is uploaded to a server. The server converts the audio data into text data and analyzes it. Based on the analysis results, feedback such as "You should practice more active listening" is generated and displayed on the smartphone.

[0527] Prompt Sentence Examples

[0528] Here are some examples of prompts the system might use:

[0529] "Please analyze the following conversation logs to extract emotional tone and areas for improvement. Additionally, please generate and provide specific feedback."

[0530] Interaction logs:

[0531] Manager: "How was work today?"

[0532] Staff: "It was hard work, but it was fun."

[0533] Feedback generated:

[0534] "The positive response was very strong. Please continue to support our staff in this way next time."

[0535] This system allows leaders and coaches to receive objective and accurate feedback, which directly contributes to improving staff skills and teaching methods.

[0536] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0537] Step 1:

[0538] Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application, which prepares the device to record the conversation. The input is the start event of the session, and the output is the status that recording is ready.

[0539] Step 2:

[0540] The device starts recording as soon as the conversation begins and captures all audio data during the conversation. The input is the audio data from the session, and the output is the recorded audio file, which is used for the subsequent analysis process.

[0541] Step 3:

[0542] After the session is over, the user stops recording and uploads the recording to the server. The input is the recorded audio file, and the output is the audio data stored on the server.

[0543] Step 4:

[0544] The server converts the received voice data into text data using a voice recognition engine. The voice recognition engine uses an algorithm to convert voice data into text information with high accuracy. The input is the uploaded voice data, and the output is the converted text data.

[0545] Step 5:

[0546] The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood. The input is text data, and the output is the analysis results of dialogue patterns and emotional tones.

[0547] Step 6:

[0548] Based on the analysis results, the server identifies areas for improvement and success and generates feedback. For example, if open dialogue is lacking or active listening is not being performed effectively, the feedback will include specific suggestions for improvement. The input is the analysis results of dialogue patterns and emotional tone, and the output is the generated feedback.

[0549] Step 7:

[0550] The generated feedback is sent from the server to the terminal, where the input is the generated feedback and the output is the feedback information presented to the user.

[0551] Step 8:

[0552] The terminal presents the received feedback to the user, who then checks the feedback and prepares for the next session. The input is the feedback information sent from the server, and the output is a feedback display that the user can check.

[0553] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0554] Embodiments of the present invention will be described in detail below.

[0555] The system of the present invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one and coaching sessions. Specifically, it solves the challenges faced by leaders and coaches through a series of processes: collecting, converting, analyzing, and recognizing emotions in audio data, as well as generating and presenting feedback.

[0556] Program flow:

[0557] Data collection phase:

[0558] 1. Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application. The user opens the application and presses the record button to prepare for recording the conversation.

[0559] 2. The device will record audio during the session in real time as instructed by the recording application. The device's microphone will capture the audio and convert it into a digital audio file. The recorded audio data will be temporarily stored on the device's storage.

[0560] 3. After the session is over, the user stops the recording and uploads the recording data to the server. The user presses the stop recording button and then clicks the "Upload Data" button, which sends the audio data from the device to the server.

[0561] Data analysis phase:

[0562] 4. The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine analyzes the voice data and starts the process of converting it into text information. The converted text data is stored on the server.

[0563] 5. The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood.

[0564] Emotion Recognition Phase:

[0565] 6. The analyzed text and voice data are then subjected to emotion analysis by the emotion engine on the server. The emotion engine recognizes emotions from the tone of voice and text and identifies the user's emotional state. This emotion data is used to generate feedback.

[0566] Feedback generation phase:

[0567] 7. Based on the analysis results and emotion recognition results, the server identifies areas for improvement and success and generates feedback. For example, the server analyzes the lack of open dialogue, lack of active listening, and appropriate emotional tone, and generates specific improvement advice as feedback. The generated feedback is stored on the server in a structured format (e.g., JSON).

[0568] 8. After the feedback is generated, the server sends it to the device. The feedback data is transferred to the device via API. The transfer method is encrypted for security reasons.

[0569] Results delivery phase:

[0570] 9. The device presents the feedback received from the server to the user. It launches a feedback display interface and visually displays areas for improvement, successes, and specific advice to the user. The user can review this information and prepare for the next session.

[0571] Examples:

[0572] For example, during an initial one-on-one session, the user begins recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results and emotion recognition by the emotion engine, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice as feedback. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[0573] This system continuously updates its algorithms based on past analysis results and the effects of feedback, improving analysis accuracy, allowing users to receive more effective feedback with repeated use, enabling leaders and coaches to continually improve their skills.

[0574] The processing flow will be explained below.

[0575] Step 1:

[0576] Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application. The user opens the application and is ready to press the record button to start the recording function.

[0577] Step 2:

[0578] The device records audio during the session in real time as directed by the recording application. The device's microphone captures the audio and stores it as a digital audio file. The recorded audio data is temporarily stored on the device's storage.

[0579] Step 3:

[0580] After the session is over, the user stops recording and uploads the recorded data to the server. The user presses the stop recording button and then clicks the "upload data" button, which sends the audio data from the device to the server.

[0581] Step 4:

[0582] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine analyzes the voice data and converts it into text information. The converted text data is stored on the server.

[0583] Step 5:

[0584] The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood. The analysis results are stored on the server and used for subsequent processing.

[0585] Step 6:

[0586] The analyzed text and voice data are then subjected to emotion analysis by an emotion engine within the server. The emotion engine recognizes emotions from voice tone and text and identifies the user's emotional state. This emotion data is then used to generate feedback.

[0587] Step 7:

[0588] Based on the analysis results and emotion recognition results, the server identifies areas for improvement and success and generates feedback. For example, the server analyzes the lack of open dialogue, lack of active listening, and appropriate emotional tone, and generates specific improvement advice as feedback. The generated feedback is stored on the server in a structured format (e.g., JSON).

[0589] Step 8:

[0590] After the feedback is generated, the server sends it to the device. The feedback data is transferred to the device via API. The transfer method uses encrypted communication to ensure security.

[0591] Step 9:

[0592] The device presents the feedback received from the server to the user, launching a feedback display interface and visually displaying areas for improvement, successes, and specific advice to the user, allowing the user to review this information and prepare for the next session.

[0593] Step 10:

[0594] Based on past analysis results and the effectiveness of feedback, the server updates the algorithm. The server evaluates past data and retrains the machine learning model to improve the system's analysis accuracy. This update occurs periodically, improving analysis accuracy and the quality of feedback.

[0595] Examples:

[0596] For example, during an initial one-on-one session, the user starts recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results and emotion recognition by the emotion engine, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice as feedback. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[0597] This system continuously updates its algorithms based on past analysis results and the effects of feedback, improving analysis accuracy, allowing users to receive more effective feedback with repeated use, enabling leaders and coaches to continually improve their skills.

[0598] Example 2

[0599] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0600] For leaders and coaches to provide effective feedback during one-on-one or coaching sessions, they need to accurately record the conversations that occur during the session and analyze the dialogue patterns and emotional tone. However, manually analyzing this information takes time and effort, and the accuracy of the analysis is often insufficient. Furthermore, if the quality of feedback does not improve, there is a problem that the skill development of leaders and coaches stagnates. The present invention aims to solve these problems and provide a system for improving the effectiveness of sessions.

[0601] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0602] In this invention, the server includes: means for a user to set up a terminal and launch a recording application before a one-on-one session or a coaching session; means for the terminal to record audio during the session in real time according to instructions from the recording application; means for the user to upload the recorded data to the server after the session ends; means for the server to pass the received audio data to a speech recognition engine and convert it into text data; means for the server to analyze the converted text data and extract dialogue patterns and emotional tones; means for an emotion engine in the server to perform emotion analysis on the analyzed text data and audio data; means for the server to generate feedback based on the analysis results and emotion recognition results; means for using a generation AI model when generating feedback; and means for presenting the generated feedback to the user. This automates the process from recording a session to analyzing it and providing feedback, making it possible to provide effective feedback quickly and accurately.

[0603] "User" refers to the person or parties who operate the recording application to collect data and review feedback when conducting one-on-one sessions or coaching sessions.

[0604] "Device" refers to the electronic device on which the user launches the recording application, collects audio data, and displays feedback, such as a smartphone or PC.

[0605] "Recording application" refers to software that uses a device to record audio from one-on-one sessions or coaching sessions in real time and save it as a digital audio file.

[0606] "Server" refers to a computer system that receives recorded voice data, analyzes the data using a voice recognition engine and emotion engine, and generates feedback.

[0607] "Speech recognition engine" refers to a technology or software module that analyzes voice data received by a server and converts it into text data.

[0608] "Text data" refers to character information converted from voice data by a voice recognition engine.

[0609] "Dialogue patterns" refer to the structure and flow of conversation extracted by analyzing text data.

[0610] "Emotional tone" refers to the emotional tendencies and tone of a speaker in a conversation.

[0611] "Emotion engine" refers to a technology or software module that performs emotion analysis based on analyzed text data and voice data to identify the user's emotional state.

[0612] "Feedback" refers to advice and specific guidelines for improvement generated based on analysis results and emotion recognition results.

[0613] A "generative AI model" is an artificial intelligence model used by the server to generate feedback, and refers to a technology that uses generative AI to generate specific improvement advice, for example.

[0614] The system of the present invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one sessions and coaching sessions. Specifically, it includes the following operations:

[0615] Data Collection Phase

[0616] First, users set up their device before a one-on-one or coaching session and launch the recording application. Users open the application on their smartphone, tablet, or computer and press the record button to begin recording the session.

[0617] The device, in accordance with the instructions of the recording application, records audio during the session in real time using the built-in microphone and converts it into a digital audio file, which is then temporarily stored in the device's storage device.

[0618] After the session ends, the user stops recording and uploads the recorded data to the server. Specifically, when the user clicks the "Stop Recording" button and then the "Upload Data" button, the audio data is sent from the device to the server using encrypted communication.

[0619] Data analysis phase

[0620] The server passes the received voice data to a speech recognition engine and converts it into text data. This speech recognition uses a common speech recognition service (e.g., Google Cloud Speech-to-Text API). The converted text data is stored in a database on the server.

[0621] The server then analyzes the converted text data and applies natural language processing (NLP) algorithms (e.g., SpaCy or NLTK) to extract dialogue patterns and emotional tones, thereby understanding the flow of dialogue and the speaker's emotional tendencies.

[0622] Emotion Recognition Phase

[0623] Based on the analyzed text and voice data, the emotion engine in the server performs emotion analysis. The emotion engine uses a deep learning model to identify the user's emotional state from the tone of voice and text. This allows it to classify emotions as positive, negative, or neutral.

[0624] Feedback generation phase

[0625] The server generates feedback based on the analysis results and emotion recognition results. This feedback is generated using a "generative AI model" (e.g., OpenAI GPT-3) to generate specific advice based on the analysis data. The generated feedback is stored on the server in a structured format (e.g., JSON).

[0626] After the feedback is generated, the server sends it to the device. The feedback data is transmitted via an API using a secure communication protocol.

[0627] Results delivery phase

[0628] The device then presents the feedback received from the server to the user, who is then presented with a dedicated application or web interface that visually displays areas for improvement, successes, and specific advice.The user can then review this information and prepare for the next session.

[0629] Specific examples

[0630] For example, during an initial one-on-one session, the user starts recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results and emotion recognition by the emotion engine, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[0631] Prompt Sentence Examples

[0632] "Please analyze the following audio data and generate feedback on problems and areas for improvement:"

[0633] "All recordings of the first 1-on-1 session"

[0634] "Analyze dialogue patterns and emotional tone"

[0635] "Generate feedback with specific advice on improving lack of open dialogue and active listening"

[0636] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0637] Step 1:

[0638] The user sets up the device and launches the recording application before the one-on-one or coaching session, which involves the user opening the dedicated recording application on their smartphone, tablet, or PC and clicking the start button.

[0639] Input: A recording application that is initiated by user action.

[0640] Output: Ready for recording.

[0641] Step 2:

[0642] The device will record audio during the session in real time as directed by the recording application, using the device's built-in microphone to convert the audio into a digital audio file that is immediately saved.

[0643] Input: Audio during the session.

[0644] Output: A digital audio file is temporarily saved to the device storage.

[0645] Step 3:

[0646] After the session ends, the user stops recording and uploads the recorded data to the server. Specifically, the user presses the application's stop recording button and then clicks the "upload data" button, which sends the audio data from the device to the server. The data is transferred using encrypted communication.

[0647] Input: Recorded audio file.

[0648] Output: The audio data is uploaded to the server.

[0649] Step 4:

[0650] The server passes the received voice data to a voice recognition engine and converts it into text data. This process includes sending the voice data to a voice recognition service (e.g., a voice recognition API) and outputting it as text data.

[0651] Input: Recorded audio data.

[0652] Output: Text data is generated and stored on the server.

[0653] Step 5:

[0654] The server analyzes the converted text data to extract dialogue patterns and emotional tones. Natural language processing (NLP) algorithms (e.g., SpaCy and NLTK) are applied to analyze the text data to identify conversation flow and emotional trends.

[0655] Input: Text data converted by speech recognition.

[0656] Output: Information on dialogue patterns and emotional tone.

[0657] Step 6:

[0658] Based on the analyzed text and voice data, the emotion engine in the server performs emotion analysis. The emotion engine uses a deep learning model to analyze the user's emotional state from voice tone and text.

[0659] Input: Information on interaction patterns and emotional tone.

[0660] Output: Sentiment classification data such as positive, negative, or neutral.

[0661] Step 7:

[0662] The server generates specific feedback based on the analysis results and emotion recognition results. This feedback is generated using a generative AI model (e.g., a generative AI model). The server creates specific improvement advice based on the analysis data and saves it in JSON format.

[0663] Input: Sentiment classification data and dialogue pattern analysis data.

[0664] Output: Feedback data is generated.

[0665] Step 8:

[0666] The server sends the generated feedback data to the device via an API using an encrypted communication protocol.

[0667] Input: Feedback data.

[0668] Output: Feedback data is sent to the terminal.

[0669] Step 9:

[0670] The device receives feedback from the server and presents it to the user. Using a dedicated application or web interface, the user is visually shown areas for improvement, successes, and specific advice. This allows the user to prepare for the next session.

[0671] Input: Feedback data.

[0672] Output: Feedback is presented to the user.

[0673] (Application example 2)

[0674] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0675] Currently, it is difficult to systematically and efficiently provide training and feedback to improve customer service skills in brick-and-mortar stores. To improve employees' customer service skills, a system is needed to evaluate their interactions with customers during their daily work and provide specific feedback on areas for improvement. Such a system must also precisely analyze emotional tone and conversation patterns to evaluate positive tone and customer service.

[0676] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0677] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for analyzing the text data and extracting dialogue patterns and emotional tones, means for evaluating customer service skills based on the analysis results and emotional tones, means for providing specific points for improvement in customer service as feedback, means for presenting the generated feedback to the user, means for updating the algorithm based on past analysis results and the effects of the feedback to improve analysis accuracy, and means for using a generative AI model to provide feedback and input prompt sentences. This makes it possible to efficiently and accurately provide specific points for improvement to improve the customer service skills of employees.

[0678] "Voice data" is information that is a digital recording of what a user says.

[0679] "Text data" is character string information obtained by analyzing voice data.

[0680] A "dialogue pattern" is information that indicates the flow and structure of a conversation, including who spoke, when, and the content of the conversation.

[0681] "Emotional tone" is information obtained by analyzing the speaker's emotional state and emotional expression during a conversation.

[0682] "Feedback" refers to specific improvements and advice regarding a user's behavior and performance that is generated based on analysis results and evaluations.

[0683] "Customer service skills" refers to the abilities and techniques that employees have when interacting with and providing service to customers in physical stores.

[0684] A "positive tone" refers to a positive and friendly manner of speaking and behavior.

[0685] "Positive customer service" refers to the act and technique of treating customers in a friendly and cheerful manner.

[0686] "Analysis precision" refers to the degree of accuracy and reproducibility of data analysis.

[0687] A "generative AI model" is a model that uses artificial intelligence technology to generate a model tailored to a specific task, and is used for data analysis and feedback generation.

[0688] A "prompt sentence" is a sentence used to input specific instructions or questions to a generative AI model.

[0689] This invention is a system for improving customer service skills in brick-and-mortar stores, and provides a series of processes for acquiring, analyzing, and generating feedback using voice data. This system has the following main hardware and software configuration:

[0690] System hardware and software configuration

[0691] 1. User device (smartphone)

[0692] Microphone: Captures audio data

[0693] Recording application: An application for recording audio data and uploading it to a server.

[0694] 2. Server

[0695] Speech recognition engine: converts voice data into text data (e.g., Google Speech Recognition)

[0696] Natural Language Processing (NLP) algorithms: Analyze text data and extract dialogue patterns and emotional tone (e.g., TextBlob)

[0697] Emotion Recognition Engine: Recognizes emotions from voice tone and text (e.g. TextBlob)

[0698] Generative AI model: Generates feedback based on analysis results

[0699] Database: Stores text data and feedback results

[0700] 3. Feedback display interface

[0701] User Interface: Present the generated feedback to the user in a GUI

[0702] Data processing and calculation

[0703] 1. Acquiring audio data

[0704] The user's smartphone records conversations with customers through a microphone.

[0705] The recording application converts the audio data into a digital format in real time and temporarily stores it in the device's storage.

[0706] 2. Uploading audio data

[0707] After finishing recording, the user clicks the "Upload Data" button and the audio data is sent to the server.

[0708] 3. Speech Recognition and Text Conversion

[0709] The server passes the received voice data to a voice recognition engine, which converts the voice data into text data.

[0710] 4. Text Data Analysis

[0711] The converted text data is then used to extract dialogue patterns and emotional tones using natural language processing (NLP) algorithms, which allow for understanding the flow of the dialogue and the speaker's emotional tendencies.

[0712] 5. Emotion analysis

[0713] The analyzed text and voice data are then subjected to emotion analysis by an emotion recognition engine, which recognizes emotions from voice tone and text to identify the user's emotional state.

[0714] 6. Feedback Generation

[0715] Based on the analysis and emotion recognition results, the server generates feedback on customer service skills and specific areas for improvement. For example, the server can determine whether there is a lack of open dialogue, a lack of active listening, or appropriate emotional tone, and provide specific advice.

[0716] 7. Providing Feedback

[0717] The generated feedback is sent from the server to the user's device and presented to the user through the device's feedback display interface, allowing the user to improve their skills based on this feedback.

[0718] Examples of concrete examples and prompts

[0719] 1. Example:

[0720] Employee A starts recording with the app while talking to Customer B.

[0721] After the conversation is over, press the "Analyze" button in the app to upload the recording data to the server.

[0722] The server converts the speech into text and performs sentiment analysis.

[0723] Based on the results of the sentiment analysis, feedback such as "Try speaking in a more positive tone" is generated.

[0724] Employee A will take that feedback into consideration at the start of their next shift.

[0725] 2. Example prompt for the generative AI model:

[0726] Prompt: Convert the following audio data into text and perform sentiment analysis. Generate appropriate feedback. Audio data: <Audio data link>

[0727] In this way, the inventive system provides specific improvements for efficiently and effectively improving employee customer service skills.

[0728] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0729] Step 1:

[0730] The user launches the recording application and presses the start recording button. The input is the user's operation, and the output is the transition to recording mode. The recording application captures the voice data during the customer interaction with the device's microphone and temporarily stores it as a digital audio file in the device's storage.

[0731] Step 2:

[0732] After finishing recording, the user presses the stop recording button. The input is the user's operation, and the output is stopping the recording. Next, the user presses the "upload data" button to send the audio data to the server. The input is the audio file, and the output is the audio file uploaded to the server.

[0733] Step 3:

[0734] The server passes the received voice data to a speech recognition engine and converts the voice file into text data. The input is the voice file, and the output is the converted text data. This process uses the Google Speech Recognition API or similar to convert voice to text.

[0735] Step 4:

[0736] The server passes the text data to a natural language processing (NLP) algorithm to extract dialogue patterns and emotional tones. The input is the text data, and the output is the analyzed dialogue patterns and emotional tones. This analysis is performed using tools such as TextBlob.

[0737] Step 5:

[0738] The server passes the analyzed text and voice data to the emotion recognition engine for emotion analysis. The input is the analyzed text and voice data, and the output is the user's emotional state data. The emotion engine analyzes the emotional tone and identifies emotions such as positive, negative, and neutral.

[0739] Step 6:

[0740] The server generates feedback using a generative AI model based on the analysis results and emotion recognition results. The inputs are dialogue patterns, emotional tone, and emotional state, and the output is specific feedback. For example, the generative AI model generates feedback such as "Try speaking in a more positive tone." The generative AI model is given instructions using prompt sentences.

[0741] Step 7:

[0742] The server sends the generated feedback to the terminal. The input is the generated feedback, and the output is the transmission of the feedback data to the terminal. The feedback data is securely transmitted using encrypted communication.

[0743] Step 8:

[0744] The terminal presents the feedback received from the server to the user. The input is feedback data, and the output is feedback displayed to the user. The feedback display interface visually presents the feedback content to the user, who can review it and use it in the next session to improve their customer service skills.

[0745] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0746] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0747] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0748] [Third embodiment]

[0749] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0750] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0751] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0752] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0753] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0754] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0755] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0756] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0757] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0758] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0759] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0760] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0761] Embodiments of the present invention will be described in detail below.

[0762] The system of this invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one and coaching sessions. Specifically, it solves the challenges faced by leaders and coaches through a series of processes: collecting, converting, and analyzing audio data, and generating and presenting feedback.

[0763] Program flow:

[0764] Data collection phase:

[0765] 1. Before the start of a 1-on-1 or coaching session, the user sets up the device and launches the recording application, which allows the entire conversation to be recorded.

[0766] 2. The device will start recording as soon as the conversation begins and will capture all audio data during the conversation.

[0767] 3. After the session is over, the user stops it and uploads the recording to the server, which provides the necessary information for the subsequent analysis process.

[0768] Data analysis phase:

[0769] 4. The server converts the received voice data into text data using a voice recognition engine. The voice recognition engine uses an algorithm to convert voice data into text information with high accuracy.

[0770] 5. The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood.

[0771] Feedback generation phase:

[0772] 6. Based on the analysis results, the server identifies areas for improvement and success and generates feedback. For example, if open dialogue is lacking or active listening is not being performed effectively, the feedback will include specific suggestions for improvement.

[0773] 7. After the feedback is generated, the server sends it to the device.

[0774] Results delivery phase:

[0775] 8. The device presents the received feedback to the user, who can then review the feedback and prepare for the next session.

[0776] Examples:

[0777] For example, during an initial one-on-one session, the user starts recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice as feedback. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[0778] This system continuously updates its algorithms based on past analysis results and the effects of feedback, improving analysis accuracy, allowing users to receive more effective feedback with repeated use, enabling leaders and coaches to continually improve their skills.

[0779] The processing flow will be explained below.

[0780] Step 1:

[0781] Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application. The user opens the application and presses the record button to prepare for recording the conversation.

[0782] Step 2:

[0783] The device records audio during the session in real time according to the instructions of the recording application. The device's microphone captures the audio and converts it into a digital audio file. The recorded audio data is temporarily stored on the device's storage.

[0784] Step 3:

[0785] After the session is over, the user stops recording and uploads the recorded data to the server. The user presses the stop recording button and then clicks the "upload data" button, which sends the audio data from the device to the server.

[0786] Step 4:

[0787] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine analyzes the voice data and starts the process of converting it into text information. The converted text data is saved on the server.

[0788] Step 5:

[0789] The server then analyzes the converted text data using natural language processing (NLP) algorithms. During this analysis process, the server identifies dialogue patterns (e.g., order and frequency of statements) and emotional tones (e.g., positive, negative) from the text data. The analysis results are stored on the server for subsequent processing.

[0790] Step 6:

[0791] The server identifies areas for improvement and success based on the analysis results and generates feedback. For example, the server detects a lack of open dialogue or active listening and generates specific improvement advice as feedback. The generated feedback is stored on the server in a structured format (e.g., JSON).

[0792] Step 7:

[0793] The server sends the generated feedback to the device. The feedback data is transferred to the device via API. The transfer method uses encrypted communication for security reasons.

[0794] Step 8:

[0795] The device presents the feedback received from the server to the user, launching a feedback display interface and visually displaying areas for improvement, successes, and specific advice to the user, allowing the user to review this information and prepare for the next session.

[0796] Step 9:

[0797] Based on past analysis results and the effectiveness of feedback, the server updates the algorithm. The server evaluates past data and retrains the machine learning model to improve the system's analysis accuracy. This update occurs periodically, improving analysis accuracy and the quality of feedback.

[0798] Example 1

[0799] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0800] Traditional leadership and coaching feedback tends to be subjective, making it difficult to provide specific and objective feedback. Furthermore, the quality of the feedback is inconsistent, making it difficult to continuously improve skills. This invention solves these problems by providing a system that provides objective and effective feedback based on voice data.

[0801] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0802] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for analyzing the text data and extracting dialogue patterns and emotional tones, means for generating feedback based on the analysis results, and means for providing the generated feedback to the user. This allows leaders and coaches to receive feedback based on specific areas for improvement and successes based on past records, enabling continuous skill development.

[0803] "Audio data" is a recording of the user's speech during a session.

[0804] "Means of collection" refers to the recording application or device used by the user before the session begins, effectively capturing the audio data.

[0805] "Means for converting into text data" refers to a voice recognition engine or software for converting collected voice data into text information.

[0806] "Means for analyzing" refers to the natural language processing (NLP) algorithms and analytical software used to extract dialogue patterns and emotional tone from text data.

[0807] "Means for generating feedback" refers to generative AI models or programs that identify areas for improvement or success based on the analysis results and provide specific advice and evaluation.

[0808] The "means for providing" refers to a communication module or application for notifying the user of the generated feedback and making the content thereof viewable.

[0809] "Open dialogue" is a method of conversation that freely draws out the opinions and thoughts of the other person and encourages active participation.

[0810] "Active listening" is a listening technique in which you listen carefully to what the other person is saying in a conversation and provide appropriate responses and feedback.

[0811] An "algorithm" refers to a series of computational steps or processing flow for solving a specific problem.

[0812] MODE FOR CARRYING OUT THE INVENTION

[0813] An embodiment of the present invention is described in detail below. The system of the present invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one and coaching sessions. Specifically, the system solves the challenges faced by leaders and coaches through a series of processes: collecting, converting, and analyzing audio data, and generating and presenting feedback.

[0814] Hardware and software configuration:

[0815] 1. Data collection phase:

[0816] Before starting a one-on-one or coaching session, the user launches a recording application on a standard device such as a smartphone, tablet, or laptop. Examples of recording applications include Otter.ai and Zoom's recording function.

[0817] By starting the recording application, the terminal enters recording mode and prepares to capture audio data during a conversation.

[0818] 2. Capture audio data:

[0819] The device automatically starts recording as soon as the conversation begins and saves all audio data, which is then continuously recorded until the session ends.

[0820] Users can simply engage in conversation without any special operations, but they can also take notes at important points.

[0821] 3. Uploading recordings:

[0822] After the session is over, the user simply closes the recording application and uploads the recording data to the server via cloud storage (e.g., Google Drive, Dropbox). This operation can also be easily performed through the application interface.

[0823] Data analysis phase:

[0824] 4. Audio to text conversion:

[0825] The server retrieves the voice data from the cloud storage and converts it into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text). This speech recognition engine converts voice into text information using a highly accurate algorithm.

[0826] 5. Text Data Analysis:

[0827] The server analyzes the generated text data using natural language processing (NLP) algorithms (e.g., SpaCy, NLTK). Specifically, it extracts dialogue patterns, analyzes emotional tones, and analyzes the frequency of key phrases and keywords. This analysis makes it possible to understand the flow of the dialogue and the speaker's emotional tendencies.

[0828] Feedback generation phase:

[0829] 6. Feedback Generation:

[0830] Based on the analysis results, the server identifies areas for improvement and success and uses a generative AI model (e.g., OpenAI GPT-3) to generate feedback. This generative model creates specific and personalized feedback based on the analysis results.

[0831] 7. Submitting Feedback:

[0832] Once the feedback is generated, the server sends it to the user's device, usually via email or a dedicated application.

[0833] Results delivery phase:

[0834] 8. Providing feedback:

[0835] The device will notify the user of the received feedback and allow them to review the details of the feedback via a dedicated app or email. Based on the feedback, users can prepare for their next session and implement specific improvements.

[0836] Examples:

[0837] For example, during an initial one-on-one session, the user launches a recording application on their smartphone and begins recording. The device starts recording at the start of the session and records the entire conversation. After the session ends, the user stops recording and uploads the data to online storage using the application's cloud storage function. The server retrieves the audio data from online storage and converts it to text using Google Cloud Speech-to-Text. The converted text data is analyzed using SpaCy to extract dialogue patterns and emotional tones. Based on this, the server generates feedback using the generative AI model GPT-3 and sends the generated feedback to the user's device. Finally, the user receives a notification, can review the feedback in the app, and prepare for the next session.

[0838] Example prompt sentence:

[0839] "I'd like some specific advice on how to improve my active listening in my next one-on-one session."

[0840] "Please give me some examples of specific words I used in past sessions to help improve our Open Dialogue."

[0841] This allows users to receive more detailed and personalized feedback and advice.

[0842] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0843] Step 1:

[0844] Before the start of a 1-on-1 or coaching session, the user launches a recording application. The user launches a dedicated recording application (e.g., general recording software) on a smartphone, tablet, or laptop. This prepares the device to record the entire conversation. The input is the user's operation, and the output is the launch of the recording application.

[0845] Step 2:

[0846] After the recording application is launched, the device enters recording mode and starts capturing audio data as soon as the conversation begins. The entire conversation is recorded and the audio data is saved in storage. The input is the user's speech, and the output is the saved audio data.

[0847] Step 3:

[0848] After the session ends, the user stops the recording application and performs an operation to upload the recorded data to cloud storage (e.g., a general cloud service). The user stops recording and then performs an operation to upload the data to the cloud. The input is stopping the recording application and the audio data, and the output is an audio file saved in cloud storage.

[0849] Step 4:

[0850] The server retrieves the uploaded audio data from the cloud storage. The server automatically accesses the cloud storage and downloads the audio data. The input is the audio file stored in the cloud storage, and the output is the audio file stored on the server.

[0851] Step 5:

[0852] The server converts the voice data into text data using a voice recognition engine (e.g., general voice recognition software). The voice recognition engine analyzes the voice data and converts it into text information. The input is voice data, and the output is the converted text data.

[0853] Step 6:

[0854] The server analyzes the converted text data using natural language processing (NLP) algorithms (e.g., general NLP libraries). Specifically, it extracts dialogue patterns, analyzes emotional tone, extracts key phrases, etc. The input is text data, and the output is the analysis results.

[0855] Step 7:

[0856] The server uses a generative AI model (e.g., a general generative model) to identify areas for improvement and success based on the analysis results and generate feedback. The generative model creates specific and personalized feedback based on the analysis results. The input is the analysis results, and the output is the generated feedback.

[0857] Step 8:

[0858] Once the feedback is generated, the server sends it to the user's device. The generated feedback is sent to the user via email or a dedicated application. The input is the feedback content, and the output is the feedback sent to the user's device.

[0859] Step 9:

[0860] The device notifies the user of the received feedback and allows them to check the details of the feedback via a dedicated application or email. The user prepares for the next session based on the feedback and implements specific improvement measures. The input is the user's feedback confirmation operation, and the output is the user's preparation for the next session.

[0861] (Application example 1)

[0862] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0863] There are limitations to the methods that leaders and coaches have for providing effective feedback to staff in one-on-one or coaching sessions, making it difficult to efficiently and effectively coach and evaluate staff, especially in physical stores. Another problem is that feedback tends to be subjective, making it difficult to accurately grasp the flow of the conversation and the emotional tone. Furthermore, the feedback generated may not directly lead to improvements in staff skills.

[0864] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0865] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for analyzing the text data and extracting dialogue patterns and emotional tones, means for generating feedback based on the analysis results, means for presenting the generated feedback to the user, means for using a natural language processing algorithm to generate the feedback, and means for using these means when training and evaluating staff in physical stores. This makes it possible to provide objective and accurate feedback based on the voice data, which can directly contribute to improving staff skills and training methods.

[0866] "Means for obtaining audio data" refers to the technical methods used to record the voices of leaders and coaches during one-on-one sessions and coaching sessions.

[0867] The "means for converting acquired voice data into text data" refers to a voice recognition technology for converting recorded voice data into text information.

[0868] "Means for analyzing text data and extracting dialogue patterns and emotional tones" refers to a technology that uses natural language processing algorithms to analyze conversation content and identify the content of statements and emotional trends.

[0869] The "means for generating feedback based on the analysis results" is a method for generating feedback based on the analyzed data that specifically indicates how the user can improve their skills and areas for improvement.

[0870] The "means for presenting the generated feedback to the user" is a technology for providing the generated feedback to the leader or coach visually or audibly.

[0871] A "natural language processing algorithm" is a computational method for understanding and analyzing text data and extracting information based on its content.

[0872] "Means of using these methods when instructing and evaluating staff in physical stores" refers to methods for leaders and coaches to use the above-mentioned technological methods when instructing and evaluating staff in physical stores.

[0873] The system of the present invention is designed to enable leaders and coaches to provide effective feedback to staff in one-on-one sessions and coaching sessions in brick-and-mortar stores. Specifically, it collects voice data, converts it into text data, and then analyzes it to generate feedback, which is then presented to the user.

[0874] Hardware and software used

[0875] The main components of this system include a smartphone, a server, a natural language processing (NLP) engine, and a voice recognition engine.

[0876] Smartphone: A device used by leaders and coaches to record and transmit audio data.

[0877] Server: A central computer system that receives and analyzes voice data.

[0878] Speech recognition engine: Software for converting voice data into text data.

[0879] Natural Language Processing Engine (NLP Engine): Provides algorithms for analyzing text data and extracting dialogue patterns and emotional tone.

[0880] Data processing and calculation

[0881] 1. The smartphone records the audio as soon as the conversation begins and uploads the conversation data to a server.

[0882] 2. The server converts the received voice data into text data using a speech recognition engine. This conversion uses a highly accurate algorithm.

[0883] 3. The converted text data is analyzed by an NLP engine on the server, which extracts dialogue patterns and emotional tones.

[0884] 4. Based on the extracted data, the server generates feedback, for example, suggesting improvements to the open dialogue if the dialogue tends to be one-way.

[0885] 5. The generated feedback is then sent back to the smartphone and presented to the leader or coach.

[0886] Specific examples

[0887] For example, consider a one-on-one session between a new staff member and the store manager at a physical store. The store manager launches an application and records the conversation. After the recording is complete, the data is uploaded to a server. The server converts the audio data into text data and analyzes it. Based on the analysis results, feedback such as "You should practice more active listening" is generated and displayed on the smartphone.

[0888] Prompt Sentence Examples

[0889] Here are some examples of prompts the system might use:

[0890] "Please analyze the following conversation logs to extract emotional tone and areas for improvement. Additionally, please generate and provide specific feedback."

[0891] Interaction logs:

[0892] Manager: "How was work today?"

[0893] Staff: "It was hard work, but it was fun."

[0894] Feedback generated:

[0895] "The positive response was very strong. Please continue to support our staff in this way next time."

[0896] This system allows leaders and coaches to receive objective and accurate feedback, which directly contributes to improving staff skills and teaching methods.

[0897] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0898] Step 1:

[0899] Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application, which prepares the device to record the conversation. The input is the start event of the session, and the output is the status that recording is ready.

[0900] Step 2:

[0901] The device starts recording as soon as the conversation begins and captures all audio data during the conversation. The input is the audio data from the session, and the output is the recorded audio file, which is used for the subsequent analysis process.

[0902] Step 3:

[0903] After the session is over, the user stops recording and uploads the recording to the server. The input is the recorded audio file, and the output is the audio data stored on the server.

[0904] Step 4:

[0905] The server converts the received voice data into text data using a voice recognition engine. The voice recognition engine uses an algorithm to convert voice data into text information with high accuracy. The input is the uploaded voice data, and the output is the converted text data.

[0906] Step 5:

[0907] The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood. The input is text data, and the output is the analysis results of dialogue patterns and emotional tones.

[0908] Step 6:

[0909] Based on the analysis results, the server identifies areas for improvement and success and generates feedback. For example, if open dialogue is lacking or active listening is not being performed effectively, the feedback will include specific suggestions for improvement. The input is the analysis results of dialogue patterns and emotional tone, and the output is the generated feedback.

[0910] Step 7:

[0911] The generated feedback is sent from the server to the terminal, where the input is the generated feedback and the output is the feedback information presented to the user.

[0912] Step 8:

[0913] The terminal presents the received feedback to the user, who then checks the feedback and prepares for the next session. The input is the feedback information sent from the server, and the output is a feedback display that the user can check.

[0914] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0915] Embodiments of the present invention will be described in detail below.

[0916] The system of the present invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one and coaching sessions. Specifically, it solves the challenges faced by leaders and coaches through a series of processes: collecting, converting, analyzing, and recognizing emotions in audio data, as well as generating and presenting feedback.

[0917] Program flow:

[0918] Data collection phase:

[0919] 1. Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application. The user opens the application and presses the record button to prepare for recording the conversation.

[0920] 2. The device will record audio during the session in real time as instructed by the recording application. The device's microphone will capture the audio and convert it into a digital audio file. The recorded audio data will be temporarily stored on the device's storage.

[0921] 3. After the session is over, the user stops the recording and uploads the recording data to the server. The user presses the stop recording button and then clicks the "Upload Data" button, which sends the audio data from the device to the server.

[0922] Data analysis phase:

[0923] 4. The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine analyzes the voice data and starts the process of converting it into text information. The converted text data is stored on the server.

[0924] 5. The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood.

[0925] Emotion Recognition Phase:

[0926] 6. The analyzed text and voice data are then subjected to emotion analysis by the emotion engine on the server. The emotion engine recognizes emotions from the tone of voice and text and identifies the user's emotional state. This emotion data is used to generate feedback.

[0927] Feedback generation phase:

[0928] 7. Based on the analysis results and emotion recognition results, the server identifies areas for improvement and success and generates feedback. For example, the server analyzes the lack of open dialogue, lack of active listening, and appropriate emotional tone, and generates specific improvement advice as feedback. The generated feedback is stored on the server in a structured format (e.g., JSON).

[0929] 8. After the feedback is generated, the server sends it to the device. The feedback data is transferred to the device via API. The transfer method is encrypted for security reasons.

[0930] Results delivery phase:

[0931] 9. The device presents the feedback received from the server to the user. It launches a feedback display interface and visually displays areas for improvement, successes, and specific advice to the user. The user can review this information and prepare for the next session.

[0932] Examples:

[0933] For example, during an initial one-on-one session, the user begins recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results and emotion recognition by the emotion engine, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice as feedback. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[0934] This system continuously updates its algorithms based on past analysis results and the effects of feedback, improving analysis accuracy, allowing users to receive more effective feedback with repeated use, enabling leaders and coaches to continually improve their skills.

[0935] The processing flow will be explained below.

[0936] Step 1:

[0937] Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application. The user opens the application and is ready to press the record button to start the recording function.

[0938] Step 2:

[0939] The device records audio during the session in real time as directed by the recording application. The device's microphone captures the audio and stores it as a digital audio file. The recorded audio data is temporarily stored on the device's storage.

[0940] Step 3:

[0941] After the session is over, the user stops recording and uploads the recorded data to the server. The user presses the stop recording button and then clicks the "upload data" button, which sends the audio data from the device to the server.

[0942] Step 4:

[0943] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine analyzes the voice data and converts it into text information. The converted text data is stored on the server.

[0944] Step 5:

[0945] The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood. The analysis results are stored on the server and used for subsequent processing.

[0946] Step 6:

[0947] The analyzed text and voice data are then subjected to emotion analysis by an emotion engine within the server. The emotion engine recognizes emotions from voice tone and text and identifies the user's emotional state. This emotion data is then used to generate feedback.

[0948] Step 7:

[0949] Based on the analysis results and emotion recognition results, the server identifies areas for improvement and success and generates feedback. For example, the server analyzes the lack of open dialogue, lack of active listening, and appropriate emotional tone, and generates specific improvement advice as feedback. The generated feedback is stored on the server in a structured format (e.g., JSON).

[0950] Step 8:

[0951] After the feedback is generated, the server sends it to the device. The feedback data is transferred to the device via API. The transfer method uses encrypted communication to ensure security.

[0952] Step 9:

[0953] The device presents the feedback received from the server to the user, launching a feedback display interface and visually displaying areas for improvement, successes, and specific advice to the user, allowing the user to review this information and prepare for the next session.

[0954] Step 10:

[0955] Based on past analysis results and the effectiveness of feedback, the server updates the algorithm. The server evaluates past data and retrains the machine learning model to improve the system's analysis accuracy. This update occurs periodically, improving analysis accuracy and the quality of feedback.

[0956] Examples:

[0957] For example, during an initial one-on-one session, the user begins recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results and emotion recognition by the emotion engine, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice as feedback. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[0958] This system continuously updates its algorithms based on past analysis results and the effects of feedback, improving analysis accuracy, allowing users to receive more effective feedback with repeated use, enabling leaders and coaches to continually improve their skills.

[0959] Example 2

[0960] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0961] For leaders and coaches to provide effective feedback during one-on-one or coaching sessions, they need to accurately record the conversations that occur during the session and analyze the dialogue patterns and emotional tone. However, manually analyzing this information takes time and effort, and the accuracy of the analysis is often insufficient. Furthermore, if the quality of feedback does not improve, there is a problem that the skill development of leaders and coaches stagnates. The present invention aims to solve these problems and provide a system for improving the effectiveness of sessions.

[0962] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0963] In this invention, the server includes: means for a user to set up a terminal and launch a recording application before a one-on-one session or a coaching session; means for the terminal to record audio during the session in real time according to instructions from the recording application; means for the user to upload the recorded data to the server after the session ends; means for the server to pass the received audio data to a speech recognition engine and convert it into text data; means for the server to analyze the converted text data and extract dialogue patterns and emotional tones; means for an emotion engine in the server to perform emotion analysis on the analyzed text data and audio data; means for the server to generate feedback based on the analysis results and emotion recognition results; means for using a generation AI model when generating feedback; and means for presenting the generated feedback to the user. This automates the process from recording a session to analyzing it and providing feedback, making it possible to provide effective feedback quickly and accurately.

[0964] "User" refers to the person or parties who operate the recording application to collect data and review feedback when conducting one-on-one sessions or coaching sessions.

[0965] "Device" refers to the electronic device on which the user launches the recording application, collects audio data, and displays feedback, such as a smartphone or PC.

[0966] "Recording application" refers to software that uses a device to record audio from one-on-one sessions or coaching sessions in real time and save it as a digital audio file.

[0967] "Server" refers to a computer system that receives recorded voice data, analyzes the data using a voice recognition engine and emotion engine, and generates feedback.

[0968] "Speech recognition engine" refers to a technology or software module that analyzes voice data received by a server and converts it into text data.

[0969] "Text data" refers to character information converted from voice data by a voice recognition engine.

[0970] "Dialogue patterns" refer to the structure and flow of conversation extracted by analyzing text data.

[0971] "Emotional tone" refers to the emotional tendencies and tone of a speaker in a conversation.

[0972] "Emotion engine" refers to a technology or software module that performs emotion analysis based on analyzed text data and voice data to identify the user's emotional state.

[0973] "Feedback" refers to advice and specific guidelines for improvement generated based on analysis results and emotion recognition results.

[0974] A "generative AI model" is an artificial intelligence model used by the server to generate feedback, and refers to a technology that uses generative AI to generate specific improvement advice, for example.

[0975] The system of the present invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one sessions and coaching sessions. Specifically, it includes the following operations:

[0976] Data Collection Phase

[0977] First, users set up their device before a one-on-one or coaching session and launch the recording application. Users open the application on their smartphone, tablet, or computer and press the record button to begin recording the session.

[0978] The device, in accordance with the instructions of the recording application, records audio during the session in real time using the built-in microphone and converts it into a digital audio file, which is then temporarily stored in the device's storage device.

[0979] After the session ends, the user stops recording and uploads the recorded data to the server. Specifically, when the user clicks the "Stop Recording" button and then the "Upload Data" button, the audio data is sent from the device to the server using encrypted communication.

[0980] Data analysis phase

[0981] The server passes the received voice data to a speech recognition engine and converts it into text data. This speech recognition uses a common speech recognition service (e.g., Google Cloud Speech-to-Text API). The converted text data is stored in a database on the server.

[0982] The server then analyzes the converted text data and applies natural language processing (NLP) algorithms (e.g., SpaCy or NLTK) to extract dialogue patterns and emotional tones, thereby understanding the flow of dialogue and the speaker's emotional tendencies.

[0983] Emotion Recognition Phase

[0984] Based on the analyzed text and voice data, the emotion engine in the server performs emotion analysis. The emotion engine uses a deep learning model to identify the user's emotional state from the tone of voice and text. This allows it to classify emotions as positive, negative, or neutral.

[0985] Feedback generation phase

[0986] The server generates feedback based on the analysis results and emotion recognition results. This feedback is generated using a "generative AI model" (e.g., OpenAI GPT-3) to generate specific advice based on the analysis data. The generated feedback is stored on the server in a structured format (e.g., JSON).

[0987] After the feedback is generated, the server sends it to the device. The feedback data is transmitted via an API using a secure communication protocol.

[0988] Results delivery phase

[0989] The device then presents the feedback received from the server to the user, who is then presented with a dedicated application or web interface that visually displays areas for improvement, successes, and specific advice.The user can then review this information and prepare for the next session.

[0990] Specific examples

[0991] For example, during an initial one-on-one session, the user starts recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results and emotion recognition by the emotion engine, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[0992] Prompt Sentence Examples

[0993] "Please analyze the following audio data and generate feedback on problems and areas for improvement:"

[0994] "All recordings of the first 1-on-1 session"

[0995] "Analyze dialogue patterns and emotional tone"

[0996] "Generate feedback with specific advice on improving lack of open dialogue and active listening"

[0997] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0998] Step 1:

[0999] The user sets up the device and launches the recording application before the one-on-one or coaching session, which involves the user opening the dedicated recording application on their smartphone, tablet, or PC and clicking the start button.

[1000] Input: A recording application that is initiated by user action.

[1001] Output: Ready for recording.

[1002] Step 2:

[1003] The device will record audio during the session in real time as directed by the recording application, using the device's built-in microphone to convert the audio into a digital audio file that is immediately saved.

[1004] Input: Audio during the session.

[1005] Output: A digital audio file is temporarily saved to the device storage.

[1006] Step 3:

[1007] After the session ends, the user stops recording and uploads the recorded data to the server. Specifically, the user presses the application's stop recording button and then clicks the "upload data" button, which sends the audio data from the device to the server. The data is transferred using encrypted communication.

[1008] Input: Recorded audio file.

[1009] Output: The audio data is uploaded to the server.

[1010] Step 4:

[1011] The server passes the received voice data to a voice recognition engine and converts it into text data. This process includes sending the voice data to a voice recognition service (e.g., a voice recognition API) and outputting it as text data.

[1012] Input: Recorded audio data.

[1013] Output: Text data is generated and stored on the server.

[1014] Step 5:

[1015] The server analyzes the converted text data to extract dialogue patterns and emotional tones. Natural language processing (NLP) algorithms (e.g., SpaCy and NLTK) are applied to analyze the text data to identify conversation flow and emotional trends.

[1016] Input: Text data converted by speech recognition.

[1017] Output: Information on dialogue patterns and emotional tone.

[1018] Step 6:

[1019] Based on the analyzed text and voice data, the emotion engine in the server performs emotion analysis. The emotion engine uses a deep learning model to analyze the user's emotional state from voice tone and text.

[1020] Input: Information on interaction patterns and emotional tone.

[1021] Output: Sentiment classification data such as positive, negative, or neutral.

[1022] Step 7:

[1023] The server generates specific feedback based on the analysis results and emotion recognition results. This feedback is generated using a generative AI model (e.g., a generative AI model). The server creates specific improvement advice based on the analysis data and saves it in JSON format.

[1024] Input: Sentiment classification data and dialogue pattern analysis data.

[1025] Output: Feedback data is generated.

[1026] Step 8:

[1027] The server sends the generated feedback data to the device via an API using an encrypted communication protocol.

[1028] Input: Feedback data.

[1029] Output: Feedback data is sent to the terminal.

[1030] Step 9:

[1031] The device receives feedback from the server and presents it to the user. Using a dedicated application or web interface, the user is visually shown areas for improvement, successes, and specific advice. This allows the user to prepare for the next session.

[1032] Input: Feedback data.

[1033] Output: Feedback is presented to the user.

[1034] (Application example 2)

[1035] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1036] Currently, it is difficult to systematically and efficiently provide training and feedback to improve customer service skills in brick-and-mortar stores. To improve employees' customer service skills, a system is needed to evaluate their interactions with customers during their daily work and provide specific feedback on areas for improvement. Such a system must also precisely analyze emotional tone and conversation patterns to evaluate positive tone and customer service.

[1037] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1038] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for analyzing the text data and extracting dialogue patterns and emotional tones, means for evaluating customer service skills based on the analysis results and emotional tones, means for providing specific points for improvement in customer service as feedback, means for presenting the generated feedback to the user, means for updating the algorithm based on past analysis results and the effects of the feedback to improve analysis accuracy, and means for using a generative AI model to provide feedback and input prompt sentences. This makes it possible to efficiently and accurately provide specific points for improvement to improve the customer service skills of employees.

[1039] "Voice data" is information that is a digital recording of what a user says.

[1040] "Text data" is character string information obtained by analyzing voice data.

[1041] A "dialogue pattern" is information that indicates the flow and structure of a conversation, including who spoke, when, and the content of the conversation.

[1042] "Emotional tone" is information obtained by analyzing the speaker's emotional state and emotional expression during a conversation.

[1043] "Feedback" refers to specific improvements and advice regarding a user's behavior and performance that is generated based on analysis results and evaluations.

[1044] "Customer service skills" refers to the abilities and techniques that employees have when interacting with and providing service to customers in physical stores.

[1045] A "positive tone" refers to a positive and friendly manner of speaking and behavior.

[1046] "Positive customer service" refers to the act and technique of treating customers in a friendly and cheerful manner.

[1047] "Analysis precision" refers to the degree of accuracy and reproducibility of data analysis.

[1048] A "generative AI model" is a model that uses artificial intelligence technology to generate a model tailored to a specific task, and is used for data analysis and feedback generation.

[1049] A "prompt sentence" is a sentence used to input specific instructions or questions to a generative AI model.

[1050] This invention is a system for improving customer service skills in brick-and-mortar stores, and provides a series of processes for acquiring, analyzing, and generating feedback using voice data. This system has the following main hardware and software configuration:

[1051] System hardware and software configuration

[1052] 1. User device (smartphone)

[1053] Microphone: Captures audio data

[1054] Recording application: An application for recording audio data and uploading it to a server.

[1055] 2. Server

[1056] Speech recognition engine: converts voice data into text data (e.g., Google Speech Recognition)

[1057] Natural Language Processing (NLP) algorithms: Analyze text data and extract dialogue patterns and emotional tone (e.g., TextBlob)

[1058] Emotion Recognition Engine: Recognizes emotions from voice tone and text (e.g. TextBlob)

[1059] Generative AI model: Generates feedback based on analysis results

[1060] Database: Stores text data and feedback results

[1061] 3. Feedback display interface

[1062] User Interface: Present the generated feedback to the user in a GUI

[1063] Data processing and calculation

[1064] 1. Acquiring audio data

[1065] The user's smartphone records conversations with customers through a microphone.

[1066] The recording application converts the audio data into a digital format in real time and temporarily stores it in the device's storage.

[1067] 2. Uploading audio data

[1068] After finishing recording, the user clicks the "Upload Data" button and the audio data is sent to the server.

[1069] 3. Speech Recognition and Text Conversion

[1070] The server passes the received voice data to a voice recognition engine, which converts the voice data into text data.

[1071] 4. Text Data Analysis

[1072] The converted text data is then used to extract dialogue patterns and emotional tones using natural language processing (NLP) algorithms, which allow for understanding the flow of the dialogue and the speaker's emotional tendencies.

[1073] 5. Emotion analysis

[1074] The analyzed text and voice data are then subjected to emotion analysis by an emotion recognition engine, which recognizes emotions from voice tone and text to identify the user's emotional state.

[1075] 6. Feedback Generation

[1076] Based on the analysis and emotion recognition results, the server generates feedback on customer service skills and specific areas for improvement. For example, the server can determine whether there is a lack of open dialogue, a lack of active listening, or appropriate emotional tone, and provide specific advice.

[1077] 7. Providing Feedback

[1078] The generated feedback is sent from the server to the user's device and presented to the user through the device's feedback display interface, allowing the user to improve their skills based on this feedback.

[1079] Examples of concrete examples and prompts

[1080] 1. Example:

[1081] Employee A starts recording with the app while talking to Customer B.

[1082] After the conversation is over, press the "Analyze" button in the app to upload the recording data to the server.

[1083] The server converts the speech into text and performs sentiment analysis.

[1084] Based on the results of the sentiment analysis, feedback such as "Try speaking in a more positive tone" is generated.

[1085] Employee A will take that feedback into consideration at the start of their next shift.

[1086] 2. Example prompt for the generative AI model:

[1087] Prompt: Convert the following audio data into text and perform sentiment analysis. Generate appropriate feedback. Audio data: <Audio data link>

[1088] In this way, the inventive system provides specific improvements for efficiently and effectively improving employee customer service skills.

[1089] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1090] Step 1:

[1091] The user launches the recording application and presses the start recording button. The input is the user's operation, and the output is the transition to recording mode. The recording application captures the voice data during the customer interaction with the device's microphone and temporarily stores it as a digital audio file in the device's storage.

[1092] Step 2:

[1093] After finishing recording, the user presses the stop recording button. The input is the user's operation, and the output is stopping the recording. Next, the user presses the "upload data" button to send the audio data to the server. The input is the audio file, and the output is the audio file uploaded to the server.

[1094] Step 3:

[1095] The server passes the received voice data to a speech recognition engine and converts the voice file into text data. The input is the voice file, and the output is the converted text data. This process uses the Google Speech Recognition API or similar to convert voice to text.

[1096] Step 4:

[1097] The server passes the text data to a natural language processing (NLP) algorithm to extract dialogue patterns and emotional tones. The input is the text data, and the output is the analyzed dialogue patterns and emotional tones. This analysis is performed using tools such as TextBlob.

[1098] Step 5:

[1099] The server passes the analyzed text and voice data to the emotion recognition engine for emotion analysis. The input is the analyzed text and voice data, and the output is the user's emotional state data. The emotion engine analyzes the emotional tone and identifies emotions such as positive, negative, and neutral.

[1100] Step 6:

[1101] The server generates feedback using a generative AI model based on the analysis results and emotion recognition results. The inputs are dialogue patterns, emotional tone, and emotional state, and the output is specific feedback. For example, the generative AI model generates feedback such as "Try speaking in a more positive tone." The generative AI model is given instructions using prompt sentences.

[1102] Step 7:

[1103] The server sends the generated feedback to the terminal. The input is the generated feedback, and the output is the transmission of the feedback data to the terminal. The feedback data is securely transmitted using encrypted communication.

[1104] Step 8:

[1105] The terminal presents the feedback received from the server to the user. The input is feedback data, and the output is feedback displayed to the user. The feedback display interface visually presents the feedback content to the user, who can review it and use it in the next session to improve their customer service skills.

[1106] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1107] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1108] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1109] [Fourth embodiment]

[1110] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1111] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1112] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1113] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1114] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1115] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1116] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1117] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1118] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1119] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1120] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1121] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1122] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1123] Embodiments of the present invention will be described in detail below.

[1124] The system of this invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one and coaching sessions. Specifically, it solves the challenges faced by leaders and coaches through a series of processes: collecting, converting, and analyzing audio data, and generating and presenting feedback.

[1125] Program flow:

[1126] Data collection phase:

[1127] 1. Before the start of a 1-on-1 or coaching session, the user sets up the device and launches the recording application, which allows the entire conversation to be recorded.

[1128] 2. The device will start recording as soon as the conversation begins and will capture all audio data during the conversation.

[1129] 3. After the session is over, the user stops it and uploads the recording to the server, which provides the necessary information for the subsequent analysis process.

[1130] Data analysis phase:

[1131] 4. The server converts the received voice data into text data using a voice recognition engine. The voice recognition engine uses an algorithm to convert voice data into text information with high accuracy.

[1132] 5. The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood.

[1133] Feedback generation phase:

[1134] 6. Based on the analysis results, the server identifies areas for improvement and success and generates feedback. For example, if open dialogue is lacking or active listening is not being performed effectively, the feedback will include specific suggestions for improvement.

[1135] 7. After the feedback is generated, the server sends it to the device.

[1136] Results delivery phase:

[1137] 8. The device presents the received feedback to the user, who can then review the feedback and prepare for the next session.

[1138] Examples:

[1139] For example, during an initial one-on-one session, the user starts recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice as feedback. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[1140] This system continuously updates its algorithms based on past analysis results and the effects of feedback, improving analysis accuracy, allowing users to receive more effective feedback with repeated use, enabling leaders and coaches to continually improve their skills.

[1141] The processing flow will be explained below.

[1142] Step 1:

[1143] Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application. The user opens the application and presses the record button to prepare for recording the conversation.

[1144] Step 2:

[1145] The device records audio during the session in real time according to the instructions of the recording application. The device's microphone captures the audio and converts it into a digital audio file. The recorded audio data is temporarily stored on the device's storage.

[1146] Step 3:

[1147] After the session is over, the user stops recording and uploads the recorded data to the server. The user presses the stop recording button and then clicks the "upload data" button, which sends the audio data from the device to the server.

[1148] Step 4:

[1149] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine analyzes the voice data and starts the process of converting it into text information. The converted text data is saved on the server.

[1150] Step 5:

[1151] The server then analyzes the converted text data using natural language processing (NLP) algorithms. During this analysis process, the server identifies dialogue patterns (e.g., order and frequency of statements) and emotional tones (e.g., positive, negative) from the text data. The analysis results are stored on the server for subsequent processing.

[1152] Step 6:

[1153] The server identifies areas for improvement and success based on the analysis results and generates feedback. For example, the server detects a lack of open dialogue or active listening and generates specific improvement advice as feedback. The generated feedback is stored on the server in a structured format (e.g., JSON).

[1154] Step 7:

[1155] The server sends the generated feedback to the device. The feedback data is transferred to the device via API. The transfer method uses encrypted communication for security reasons.

[1156] Step 8:

[1157] The device presents the feedback received from the server to the user, launching a feedback display interface and visually displaying areas for improvement, successes, and specific advice to the user, allowing the user to review this information and prepare for the next session.

[1158] Step 9:

[1159] Based on past analysis results and the effectiveness of feedback, the server updates the algorithm. The server evaluates past data and retrains the machine learning model to improve the system's analysis accuracy. This update occurs periodically, improving analysis accuracy and the quality of feedback.

[1160] Example 1

[1161] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1162] Traditional leadership and coaching feedback tends to be subjective, making it difficult to provide specific and objective feedback. Furthermore, the quality of the feedback is inconsistent, making it difficult to continuously improve skills. This invention solves these problems by providing a system that provides objective and effective feedback based on voice data.

[1163] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1164] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for analyzing the text data and extracting dialogue patterns and emotional tones, means for generating feedback based on the analysis results, and means for providing the generated feedback to the user. This allows leaders and coaches to receive feedback based on specific areas for improvement and successes based on past records, enabling continuous skill development.

[1165] "Audio data" is a recording of the user's speech during a session.

[1166] "Means of collection" refers to the recording application or device used by the user before the session begins, effectively capturing the audio data.

[1167] "Means for converting into text data" refers to a voice recognition engine or software for converting collected voice data into text information.

[1168] "Means for analyzing" refers to the natural language processing (NLP) algorithms and analytical software used to extract dialogue patterns and emotional tone from text data.

[1169] "Means for generating feedback" refers to generative AI models or programs that identify areas for improvement or success based on the analysis results and provide specific advice and evaluation.

[1170] The "means for providing" refers to a communication module or application for notifying the user of the generated feedback and making the content thereof viewable.

[1171] "Open dialogue" is a method of conversation that freely draws out the opinions and thoughts of the other person and encourages active participation.

[1172] "Active listening" is a listening technique in which you listen carefully to what the other person is saying in a conversation and provide appropriate responses and feedback.

[1173] An "algorithm" refers to a series of computational steps or processing flow for solving a specific problem.

[1174] MODE FOR CARRYING OUT THE INVENTION

[1175] An embodiment of the present invention is described in detail below. The system of the present invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one and coaching sessions. Specifically, the system solves the challenges faced by leaders and coaches through a series of processes: collecting, converting, and analyzing audio data, and generating and presenting feedback.

[1176] Hardware and software configuration:

[1177] 1. Data collection phase:

[1178] Before starting a one-on-one or coaching session, the user launches a recording application on a standard device such as a smartphone, tablet, or laptop. Examples of recording applications include Otter.ai and Zoom's recording function.

[1179] By starting the recording application, the terminal enters recording mode and prepares to capture audio data during a conversation.

[1180] 2. Capture audio data:

[1181] The device automatically starts recording as soon as the conversation begins and saves all audio data, which is then continuously recorded until the session ends.

[1182] Users can simply engage in conversation without any special operations, but they can also take notes at important points.

[1183] 3. Uploading recordings:

[1184] After the session is over, the user simply closes the recording application and uploads the recording data to the server via cloud storage (e.g., Google Drive, Dropbox). This operation can also be easily performed through the application interface.

[1185] Data analysis phase:

[1186] 4. Audio to text conversion:

[1187] The server retrieves the voice data from the cloud storage and converts it into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text). This speech recognition engine converts voice into text information using a highly accurate algorithm.

[1188] 5. Text Data Analysis:

[1189] The server analyzes the generated text data using natural language processing (NLP) algorithms (e.g., SpaCy, NLTK). Specifically, it extracts dialogue patterns, analyzes emotional tones, and analyzes the frequency of key phrases and keywords. This analysis makes it possible to understand the flow of the dialogue and the speaker's emotional tendencies.

[1190] Feedback generation phase:

[1191] 6. Feedback Generation:

[1192] Based on the analysis results, the server identifies areas for improvement and success and uses a generative AI model (e.g., OpenAI GPT-3) to generate feedback. This generative model creates specific and personalized feedback based on the analysis results.

[1193] 7. Submitting Feedback:

[1194] Once the feedback is generated, the server sends it to the user's device, usually via email or a dedicated application.

[1195] Results delivery phase:

[1196] 8. Providing feedback:

[1197] The device will notify the user of the received feedback and allow them to review the details of the feedback via a dedicated app or email. Based on the feedback, users can prepare for their next session and implement specific improvements.

[1198] Examples:

[1199] For example, during an initial one-on-one session, the user launches a recording application on their smartphone and begins recording. The device starts recording at the start of the session and records the entire conversation. After the session ends, the user stops recording and uploads the data to online storage using the application's cloud storage function. The server retrieves the audio data from online storage and converts it to text using Google Cloud Speech-to-Text. The converted text data is analyzed using SpaCy to extract dialogue patterns and emotional tones. Based on this, the server generates feedback using the generative AI model GPT-3 and sends the generated feedback to the user's device. Finally, the user receives a notification, can review the feedback in the app, and prepare for the next session.

[1200] Example prompt sentence:

[1201] "I'd like some specific advice on how to improve my active listening in my next one-on-one session."

[1202] "Please give me some examples of specific words I used in past sessions to help improve our Open Dialogue."

[1203] This allows users to receive more detailed and personalized feedback and advice.

[1204] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1205] Step 1:

[1206] Before the start of a 1-on-1 or coaching session, the user launches a recording application. The user launches a dedicated recording application (e.g., general recording software) on a smartphone, tablet, or laptop. This prepares the device to record the entire conversation. The input is the user's operation, and the output is the launch of the recording application.

[1207] Step 2:

[1208] After the recording application is launched, the device enters recording mode and starts capturing audio data as soon as the conversation begins. The entire conversation is recorded and the audio data is saved in storage. The input is the user's speech, and the output is the saved audio data.

[1209] Step 3:

[1210] After the session ends, the user stops the recording application and performs an operation to upload the recorded data to cloud storage (e.g., a general cloud service). The user stops recording and then performs an operation to upload the data to the cloud. The input is stopping the recording application and the audio data, and the output is an audio file saved in cloud storage.

[1211] Step 4:

[1212] The server retrieves the uploaded audio data from the cloud storage. The server automatically accesses the cloud storage and downloads the audio data. The input is the audio file stored in the cloud storage, and the output is the audio file stored on the server.

[1213] Step 5:

[1214] The server converts the voice data into text data using a voice recognition engine (e.g., general voice recognition software). The voice recognition engine analyzes the voice data and converts it into text information. The input is voice data, and the output is the converted text data.

[1215] Step 6:

[1216] The server analyzes the converted text data using natural language processing (NLP) algorithms (e.g., general NLP libraries). Specifically, it extracts dialogue patterns, analyzes emotional tone, extracts key phrases, etc. The input is text data, and the output is the analysis results.

[1217] Step 7:

[1218] The server uses a generative AI model (e.g., a general generative model) to identify areas for improvement and success based on the analysis results and generate feedback. The generative model creates specific and personalized feedback based on the analysis results. The input is the analysis results, and the output is the generated feedback.

[1219] Step 8:

[1220] Once the feedback is generated, the server sends it to the user's device. The generated feedback is sent to the user via email or a dedicated application. The input is the feedback content, and the output is the feedback sent to the user's device.

[1221] Step 9:

[1222] The device notifies the user of the received feedback and allows them to check the details of the feedback via a dedicated application or email. The user prepares for the next session based on the feedback and implements specific improvement measures. The input is the user's feedback confirmation operation, and the output is the user's preparation for the next session.

[1223] (Application example 1)

[1224] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1225] There are limitations to the methods that leaders and coaches have for providing effective feedback to staff in one-on-one or coaching sessions, making it difficult to efficiently and effectively coach and evaluate staff, especially in physical stores. Another problem is that feedback tends to be subjective, making it difficult to accurately grasp the flow of the conversation and the emotional tone. Furthermore, the feedback generated may not directly lead to improvements in staff skills.

[1226] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1227] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for analyzing the text data and extracting dialogue patterns and emotional tones, means for generating feedback based on the analysis results, means for presenting the generated feedback to the user, means for using a natural language processing algorithm to generate the feedback, and means for using these means when training and evaluating staff in physical stores. This makes it possible to provide objective and accurate feedback based on the voice data, which can directly contribute to improving staff skills and training methods.

[1228] "Means for obtaining audio data" refers to the technical methods used to record the voices of leaders and coaches during one-on-one sessions and coaching sessions.

[1229] The "means for converting acquired voice data into text data" refers to a voice recognition technology for converting recorded voice data into text information.

[1230] "Means for analyzing text data and extracting dialogue patterns and emotional tones" refers to a technology that uses natural language processing algorithms to analyze conversation content and identify the content of statements and emotional trends.

[1231] The "means for generating feedback based on the analysis results" is a method for generating feedback based on the analyzed data that specifically indicates how the user can improve their skills and areas for improvement.

[1232] The "means for presenting the generated feedback to the user" is a technology for providing the generated feedback to the leader or coach visually or audibly.

[1233] A "natural language processing algorithm" is a computational method for understanding and analyzing text data and extracting information based on its content.

[1234] "Means of using these methods when instructing and evaluating staff in physical stores" refers to methods for leaders and coaches to use the above-mentioned technological methods when instructing and evaluating staff in physical stores.

[1235] The system of the present invention is designed to enable leaders and coaches to provide effective feedback to staff in one-on-one sessions and coaching sessions in brick-and-mortar stores. Specifically, it collects voice data, converts it into text data, and then analyzes it to generate feedback, which is then presented to the user.

[1236] Hardware and software used

[1237] The main components of this system include a smartphone, a server, a natural language processing (NLP) engine, and a voice recognition engine.

[1238] Smartphone: A device used by leaders and coaches to record and transmit audio data.

[1239] Server: A central computer system that receives and analyzes voice data.

[1240] Speech recognition engine: Software for converting voice data into text data.

[1241] Natural Language Processing Engine (NLP Engine): Provides algorithms for analyzing text data and extracting dialogue patterns and emotional tone.

[1242] Data processing and calculation

[1243] 1. The smartphone records the audio as soon as the conversation begins and uploads the conversation data to a server.

[1244] 2. The server converts the received voice data into text data using a speech recognition engine. This conversion uses a highly accurate algorithm.

[1245] 3. The converted text data is analyzed by an NLP engine on the server, which extracts dialogue patterns and emotional tones.

[1246] 4. Based on the extracted data, the server generates feedback, for example, suggesting improvements to the open dialogue if the dialogue tends to be one-way.

[1247] 5. The generated feedback is then sent back to the smartphone and presented to the leader or coach.

[1248] Specific examples

[1249] For example, consider a one-on-one session between a new staff member and the store manager at a physical store. The store manager launches an application and records the conversation. After the recording is complete, the data is uploaded to a server. The server converts the audio data into text data and analyzes it. Based on the analysis results, feedback such as "You should practice more active listening" is generated and displayed on the smartphone.

[1250] Prompt Sentence Examples

[1251] Here are some examples of prompts the system might use:

[1252] "Please analyze the following conversation logs to extract emotional tone and areas for improvement. Additionally, please generate and provide specific feedback."

[1253] Interaction logs:

[1254] Manager: "How was work today?"

[1255] Staff: "It was hard work, but it was fun."

[1256] Feedback generated:

[1257] "The positive response was very strong. Please continue to support our staff in this way next time."

[1258] This system allows leaders and coaches to receive objective and accurate feedback, which directly contributes to improving staff skills and teaching methods.

[1259] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1260] Step 1:

[1261] Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application, which prepares the device to record the conversation. The input is the start event of the session, and the output is the status that recording is ready.

[1262] Step 2:

[1263] The device starts recording as soon as the conversation begins and captures all audio data during the conversation. The input is the audio data from the session, and the output is the recorded audio file, which is used for the subsequent analysis process.

[1264] Step 3:

[1265] After the session is over, the user stops recording and uploads the recording to the server. The input is the recorded audio file, and the output is the audio data stored on the server.

[1266] Step 4:

[1267] The server converts the received voice data into text data using a voice recognition engine. The voice recognition engine uses an algorithm to convert voice data into text information with high accuracy. The input is the uploaded voice data, and the output is the converted text data.

[1268] Step 5:

[1269] The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood. The input is text data, and the output is the analysis results of dialogue patterns and emotional tones.

[1270] Step 6:

[1271] Based on the analysis results, the server identifies areas for improvement and success and generates feedback. For example, if open dialogue is lacking or active listening is not being performed effectively, the feedback will include specific suggestions for improvement. The input is the analysis results of dialogue patterns and emotional tone, and the output is the generated feedback.

[1272] Step 7:

[1273] The generated feedback is sent from the server to the terminal, where the input is the generated feedback and the output is the feedback information presented to the user.

[1274] Step 8:

[1275] The terminal presents the received feedback to the user, who then checks the feedback and prepares for the next session. The input is the feedback information sent from the server, and the output is a feedback display that the user can check.

[1276] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1277] Embodiments of the present invention will be described in detail below.

[1278] The system of the present invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one and coaching sessions. Specifically, it solves the challenges faced by leaders and coaches through a series of processes: collecting, converting, analyzing, and recognizing emotions in audio data, as well as generating and presenting feedback.

[1279] Program flow:

[1280] Data collection phase:

[1281] 1. Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application. The user opens the application and presses the record button to prepare for recording the conversation.

[1282] 2. The device will record audio during the session in real time as instructed by the recording application. The device's microphone will capture the audio and convert it into a digital audio file. The recorded audio data will be temporarily stored on the device's storage.

[1283] 3. After the session is over, the user stops the recording and uploads the recording data to the server. The user presses the stop recording button and then clicks the "Upload Data" button, which sends the audio data from the device to the server.

[1284] Data analysis phase:

[1285] 4. The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine analyzes the voice data and starts the process of converting it into text information. The converted text data is stored on the server.

[1286] 5. The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood.

[1287] Emotion Recognition Phase:

[1288] 6. The analyzed text and voice data are then subjected to emotion analysis by the emotion engine on the server. The emotion engine recognizes emotions from the tone of voice and text and identifies the user's emotional state. This emotion data is used to generate feedback.

[1289] Feedback generation phase:

[1290] 7. Based on the analysis results and emotion recognition results, the server identifies areas for improvement and success and generates feedback. For example, the server analyzes the lack of open dialogue, lack of active listening, and appropriate emotional tone, and generates specific improvement advice as feedback. The generated feedback is stored on the server in a structured format (e.g., JSON).

[1291] 8. After the feedback is generated, the server sends it to the device. The feedback data is transferred to the device via API. The transfer method is encrypted for security reasons.

[1292] Results delivery phase:

[1293] 9. The device presents the feedback received from the server to the user. It launches a feedback display interface and visually displays areas for improvement, successes, and specific advice to the user. The user can review this information and prepare for the next session.

[1294] Examples:

[1295] For example, during an initial one-on-one session, the user begins recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results and emotion recognition by the emotion engine, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice as feedback. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[1296] This system continuously updates its algorithms based on past analysis results and the effects of feedback, improving analysis accuracy, allowing users to receive more effective feedback with repeated use, enabling leaders and coaches to continually improve their skills.

[1297] The processing flow will be explained below.

[1298] Step 1:

[1299] Before starting a 1-on-1 or coaching session, the user sets up the device and launches the recording application. The user opens the application and is ready to press the record button to start the recording function.

[1300] Step 2:

[1301] The device records audio during the session in real time as directed by the recording application. The device's microphone captures the audio and stores it as a digital audio file. The recorded audio data is temporarily stored on the device's storage.

[1302] Step 3:

[1303] After the session is over, the user stops recording and uploads the recorded data to the server. The user presses the stop recording button and then clicks the "upload data" button, which sends the audio data from the device to the server.

[1304] Step 4:

[1305] The server passes the received voice data to a voice recognition engine and converts it into text data. The voice recognition engine analyzes the voice data and converts it into text information. The converted text data is stored on the server.

[1306] Step 5:

[1307] The converted text data is then analyzed by the server. Specifically, natural language processing (NLP) algorithms are used to extract dialogue patterns and emotional tones. This analysis allows the flow of the dialogue and the speaker's emotional tendencies to be understood. The analysis results are stored on the server and used for subsequent processing.

[1308] Step 6:

[1309] The analyzed text and voice data are then subjected to emotion analysis by an emotion engine within the server. The emotion engine recognizes emotions from voice tone and text and identifies the user's emotional state. This emotion data is then used to generate feedback.

[1310] Step 7:

[1311] Based on the analysis results and emotion recognition results, the server identifies areas for improvement and success and generates feedback. For example, the server analyzes the lack of open dialogue, lack of active listening, and appropriate emotional tone, and generates specific improvement advice as feedback. The generated feedback is stored on the server in a structured format (e.g., JSON).

[1312] Step 8:

[1313] After the feedback is generated, the server sends it to the device. The feedback data is transferred to the device via API. The transfer method uses encrypted communication to ensure security.

[1314] Step 9:

[1315] The device presents the feedback received from the server to the user, launching a feedback display interface and visually displaying areas for improvement, successes, and specific advice to the user, allowing the user to review this information and prepare for the next session.

[1316] Step 10:

[1317] Based on past analysis results and the effectiveness of feedback, the server updates the algorithm. The server evaluates past data and retrains the machine learning model to improve the system's analysis accuracy. This update occurs periodically, improving analysis accuracy and the quality of feedback.

[1318] Examples:

[1319] For example, during an initial one-on-one session, the user begins recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results and emotion recognition by the emotion engine, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice as feedback. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[1320] This system continuously updates its algorithms based on past analysis results and the effects of feedback, improving analysis accuracy, allowing users to receive more effective feedback with repeated use, enabling leaders and coaches to continually improve their skills.

[1321] Example 2

[1322] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1323] For leaders and coaches to provide effective feedback during one-on-one or coaching sessions, they need to accurately record the conversations that occur during the session and analyze the dialogue patterns and emotional tone. However, manually analyzing this information takes time and effort, and the accuracy of the analysis is often insufficient. Furthermore, if the quality of feedback does not improve, there is a problem that the skill development of leaders and coaches stagnates. The present invention aims to solve these problems and provide a system for improving the effectiveness of sessions.

[1324] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1325] In this invention, the server includes: means for a user to set up a terminal and launch a recording application before a one-on-one session or a coaching session; means for the terminal to record audio during the session in real time according to instructions from the recording application; means for the user to upload the recorded data to the server after the session ends; means for the server to pass the received audio data to a speech recognition engine and convert it into text data; means for the server to analyze the converted text data and extract dialogue patterns and emotional tones; means for an emotion engine in the server to perform emotion analysis on the analyzed text data and audio data; means for the server to generate feedback based on the analysis results and emotion recognition results; means for using a generation AI model when generating feedback; and means for presenting the generated feedback to the user. This automates the process from recording a session to analyzing it and providing feedback, making it possible to provide effective feedback quickly and accurately.

[1326] "User" refers to the person or parties who operate the recording application to collect data and review feedback when conducting one-on-one sessions or coaching sessions.

[1327] "Device" refers to the electronic device on which the user launches the recording application, collects audio data, and displays feedback, such as a smartphone or PC.

[1328] "Recording application" refers to software that uses a device to record audio from one-on-one sessions or coaching sessions in real time and save it as a digital audio file.

[1329] "Server" refers to a computer system that receives recorded voice data, analyzes the data using a voice recognition engine and emotion engine, and generates feedback.

[1330] "Speech recognition engine" refers to a technology or software module that analyzes voice data received by a server and converts it into text data.

[1331] "Text data" refers to character information converted from voice data by a voice recognition engine.

[1332] "Dialogue patterns" refer to the structure and flow of conversation extracted by analyzing text data.

[1333] "Emotional tone" refers to the emotional tendencies and tone of a speaker in a conversation.

[1334] "Emotion engine" refers to a technology or software module that performs emotion analysis based on analyzed text data and voice data to identify the user's emotional state.

[1335] "Feedback" refers to advice and specific guidelines for improvement generated based on analysis results and emotion recognition results.

[1336] A "generative AI model" is an artificial intelligence model used by the server to generate feedback, and refers to a technology that uses generative AI to generate specific improvement advice, for example.

[1337] The system of the present invention is designed to help leaders and coaches improve their skills and provide effective feedback in one-on-one sessions and coaching sessions. Specifically, it includes the following operations:

[1338] Data Collection Phase

[1339] First, users set up their device before a one-on-one or coaching session and launch the recording application. Users open the application on their smartphone, tablet, or computer and press the record button to begin recording the session.

[1340] The device, in accordance with the instructions of the recording application, records audio during the session in real time using the built-in microphone and converts it into a digital audio file, which is then temporarily stored in the device's storage device.

[1341] After the session ends, the user stops recording and uploads the recorded data to the server. Specifically, when the user clicks the "Stop Recording" button and then the "Upload Data" button, the audio data is sent from the device to the server using encrypted communication.

[1342] Data analysis phase

[1343] The server passes the received voice data to a speech recognition engine and converts it into text data. This speech recognition uses a common speech recognition service (e.g., Google Cloud Speech-to-Text API). The converted text data is stored in a database on the server.

[1344] The server then analyzes the converted text data and applies natural language processing (NLP) algorithms (e.g., SpaCy or NLTK) to extract dialogue patterns and emotional tones, thereby understanding the flow of dialogue and the speaker's emotional tendencies.

[1345] Emotion Recognition Phase

[1346] Based on the analyzed text and voice data, the emotion engine in the server performs emotion analysis. The emotion engine uses a deep learning model to identify the user's emotional state from the tone of voice and text. This allows it to classify emotions as positive, negative, or neutral.

[1347] Feedback generation phase

[1348] The server generates feedback based on the analysis results and emotion recognition results. This feedback is generated using a "generative AI model" (e.g., OpenAI GPT-3) to generate specific advice based on the analysis data. The generated feedback is stored on the server in a structured format (e.g., JSON).

[1349] After the feedback is generated, the server sends it to the device. The feedback data is transmitted via an API using a secure communication protocol.

[1350] Results delivery phase

[1351] The device then presents the feedback received from the server to the user, who is then presented with a dedicated application or web interface that visually displays areas for improvement, successes, and specific advice.The user can then review this information and prepare for the next session.

[1352] Specific examples

[1353] For example, during an initial one-on-one session, the user starts recording the conversation. The device records the conversation and sends the data to the server after the session ends. The server converts the voice data into text and analyzes the dialogue patterns and emotional tone. Based on the analysis results and emotion recognition by the emotion engine, the server determines that improvements to open dialogue and active listening are necessary and generates specific advice. The generated feedback is sent to the device, where the user can review it and implement improvements for the next session.

[1354] Prompt Sentence Examples

[1355] "Please analyze the following audio data and generate feedback on problems and areas for improvement:"

[1356] "All recordings of the first 1-on-1 session"

[1357] "Analyze dialogue patterns and emotional tone"

[1358] "Generate feedback with specific advice on improving lack of open dialogue and active listening"

[1359] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1360] Step 1:

[1361] The user sets up the device and launches the recording application before the one-on-one or coaching session, which involves the user opening the dedicated recording application on their smartphone, tablet, or PC and clicking the start button.

[1362] Input: A recording application that is initiated by user action.

[1363] Output: Ready for recording.

[1364] Step 2:

[1365] The device will record audio during the session in real time as directed by the recording application, using the device's built-in microphone to convert the audio into a digital audio file that is immediately saved.

[1366] Input: Audio during the session.

[1367] Output: A digital audio file is temporarily saved to the device storage.

[1368] Step 3:

[1369] After the session ends, the user stops recording and uploads the recorded data to the server. Specifically, the user presses the application's stop recording button and then clicks the "upload data" button, which sends the audio data from the device to the server. The data is transferred using encrypted communication.

[1370] Input: Recorded audio file.

[1371] Output: The audio data is uploaded to the server.

[1372] Step 4:

[1373] The server passes the received voice data to a voice recognition engine and converts it into text data. This process includes sending the voice data to a voice recognition service (e.g., a voice recognition API) and outputting it as text data.

[1374] Input: Recorded audio data.

[1375] Output: Text data is generated and stored on the server.

[1376] Step 5:

[1377] The server analyzes the converted text data to extract dialogue patterns and emotional tones. Natural language processing (NLP) algorithms (e.g., SpaCy and NLTK) are applied to analyze the text data to identify conversation flow and emotional trends.

[1378] Input: Text data converted by speech recognition.

[1379] Output: Information on dialogue patterns and emotional tone.

[1380] Step 6:

[1381] Based on the analyzed text and voice data, the emotion engine in the server performs emotion analysis. The emotion engine uses a deep learning model to analyze the user's emotional state from voice tone and text.

[1382] Input: Information on interaction patterns and emotional tone.

[1383] Output: Sentiment classification data such as positive, negative, or neutral.

[1384] Step 7:

[1385] The server generates specific feedback based on the analysis results and emotion recognition results. This feedback is generated using a generative AI model (e.g., a generative AI model). The server creates specific improvement advice based on the analysis data and saves it in JSON format.

[1386] Input: Sentiment classification data and dialogue pattern analysis data.

[1387] Output: Feedback data is generated.

[1388] Step 8:

[1389] The server sends the generated feedback data to the device via an API using an encrypted communication protocol.

[1390] Input: Feedback data.

[1391] Output: Feedback data is sent to the terminal.

[1392] Step 9:

[1393] The device receives feedback from the server and presents it to the user. Using a dedicated application or web interface, the user is visually shown areas for improvement, successes, and specific advice. This allows the user to prepare for the next session.

[1394] Input: Feedback data.

[1395] Output: Feedback is presented to the user.

[1396] (Application example 2)

[1397] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1398] Currently, it is difficult to systematically and efficiently provide training and feedback to improve customer service skills in brick-and-mortar stores. To improve employees' customer service skills, a system is needed to evaluate their interactions with customers during their daily work and provide specific feedback on areas for improvement. Such a system must also precisely analyze emotional tone and conversation patterns to evaluate positive tone and customer service.

[1399] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1400] In this invention, the server includes means for acquiring voice data, means for converting the acquired voice data into text data, means for analyzing the text data and extracting dialogue patterns and emotional tones, means for evaluating customer service skills based on the analysis results and emotional tones, means for providing specific points for improvement in customer service as feedback, means for presenting the generated feedback to the user, means for updating the algorithm based on past analysis results and the effects of the feedback to improve analysis accuracy, and means for using a generative AI model to provide feedback and input prompt sentences. This makes it possible to efficiently and accurately provide specific points for improvement to improve the customer service skills of employees.

[1401] "Voice data" is information that is a digital recording of what a user says.

[1402] "Text data" is character string information obtained by analyzing voice data.

[1403] A "dialogue pattern" is information that indicates the flow and structure of a conversation, including who spoke, when, and the content of the conversation.

[1404] "Emotional tone" is information obtained by analyzing the speaker's emotional state and emotional expression during a conversation.

[1405] "Feedback" refers to specific improvements and advice regarding a user's behavior and performance that is generated based on analysis results and evaluations.

[1406] "Customer service skills" refers to the abilities and techniques that employees have when interacting with and providing service to customers in physical stores.

[1407] A "positive tone" refers to a positive and friendly manner of speaking and behavior.

[1408] "Positive customer service" refers to the act and technique of treating customers in a friendly and cheerful manner.

[1409] "Analysis precision" refers to the degree of accuracy and reproducibility of data analysis.

[1410] A "generative AI model" is a model that uses artificial intelligence technology to generate a model tailored to a specific task, and is used for data analysis and feedback generation.

[1411] A "prompt sentence" is a sentence used to input specific instructions or questions to a generative AI model.

[1412] This invention is a system for improving customer service skills in brick-and-mortar stores, and provides a series of processes for acquiring, analyzing, and generating feedback using voice data. This system has the following main hardware and software configuration:

[1413] System hardware and software configuration

[1414] 1. User device (smartphone)

[1415] Microphone: Captures audio data

[1416] Recording application: An application for recording audio data and uploading it to a server.

[1417] 2. Server

[1418] Speech recognition engine: converts voice data into text data (e.g., Google Speech Recognition)

[1419] Natural Language Processing (NLP) algorithms: Analyze text data and extract dialogue patterns and emotional tone (e.g., TextBlob)

[1420] Emotion Recognition Engine: Recognizes emotions from voice tone and text (e.g. TextBlob)

[1421] Generative AI model: Generates feedback based on analysis results

[1422] Database: Stores text data and feedback results

[1423] 3. Feedback display interface

[1424] User Interface: Present the generated feedback to the user in a GUI

[1425] Data processing and calculation

[1426] 1. Acquiring audio data

[1427] The user's smartphone records conversations with customers through a microphone.

[1428] The recording application converts the audio data into a digital format in real time and temporarily stores it in the device's storage.

[1429] 2. Uploading audio data

[1430] After finishing recording, the user clicks the "Upload Data" button and the audio data is sent to the server.

[1431] 3. Speech Recognition and Text Conversion

[1432] The server passes the received voice data to a voice recognition engine, which converts the voice data into text data.

[1433] 4. Text Data Analysis

[1434] The converted text data is then used to extract dialogue patterns and emotional tones using natural language processing (NLP) algorithms, which allow for understanding the flow of the dialogue and the speaker's emotional tendencies.

[1435] 5. Emotion analysis

[1436] The analyzed text and voice data are then subjected to emotion analysis by an emotion recognition engine, which recognizes emotions from voice tone and text to identify the user's emotional state.

[1437] 6. Feedback Generation

[1438] Based on the analysis and emotion recognition results, the server generates feedback on customer service skills and specific areas for improvement. For example, the server can determine whether there is a lack of open dialogue, a lack of active listening, or appropriate emotional tone, and provide specific advice.

[1439] 7. Providing Feedback

[1440] The generated feedback is sent from the server to the user's device and presented to the user through the device's feedback display interface, allowing the user to improve their skills based on this feedback.

[1441] Examples of concrete examples and prompts

[1442] 1. Example:

[1443] Employee A starts recording with the app while talking to Customer B.

[1444] After the conversation is over, press the "Analyze" button in the app to upload the recording data to the server.

[1445] The server converts the speech into text and performs sentiment analysis.

[1446] Based on the results of the sentiment analysis, feedback such as "Try speaking in a more positive tone" is generated.

[1447] Employee A will take that feedback into consideration at the start of their next shift.

[1448] 2. Example prompt for the generative AI model:

[1449] Prompt: Convert the following audio data into text and perform sentiment analysis. Generate appropriate feedback. Audio data: <Audio data link>

[1450] In this way, the inventive system provides specific improvements for efficiently and effectively improving employee customer service skills.

[1451] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1452] Step 1:

[1453] The user launches the recording application and presses the start recording button. The input is the user's operation, and the output is the transition to recording mode. The recording application captures the voice data during the customer interaction with the device's microphone and temporarily stores it as a digital audio file in the device's storage.

[1454] Step 2:

[1455] After finishing recording, the user presses the stop recording button. The input is the user's operation, and the output is stopping the recording. Next, the user presses the "upload data" button to send the audio data to the server. The input is the audio file, and the output is the audio file uploaded to the server.

[1456] Step 3:

[1457] The server passes the received voice data to a speech recognition engine and converts the voice file into text data. The input is the voice file, and the output is the converted text data. This process uses the Google Speech Recognition API or similar to convert voice to text.

[1458] Step 4:

[1459] The server passes the text data to a natural language processing (NLP) algorithm to extract dialogue patterns and emotional tones. The input is the text data, and the output is the analyzed dialogue patterns and emotional tones. This analysis is performed using tools such as TextBlob.

[1460] Step 5:

[1461] The server passes the analyzed text and voice data to the emotion recognition engine for emotion analysis. The input is the analyzed text and voice data, and the output is the user's emotional state data. The emotion engine analyzes the emotional tone and identifies emotions such as positive, negative, and neutral.

[1462] Step 6:

[1463] The server generates feedback using a generative AI model based on the analysis results and emotion recognition results. The inputs are dialogue patterns, emotional tone, and emotional state, and the output is specific feedback. For example, the generative AI model generates feedback such as "Try speaking in a more positive tone." The generative AI model is given instructions using prompt sentences.

[1464] Step 7:

[1465] The server sends the generated feedback to the terminal. The input is the generated feedback, and the output is the transmission of the feedback data to the terminal. The feedback data is securely transmitted using encrypted communication.

[1466] Step 8:

[1467] The terminal presents the feedback received from the server to the user. The input is feedback data, and the output is feedback displayed to the user. The feedback display interface visually presents the feedback content to the user, who can review it and use it in the next session to improve their customer service skills.

[1468] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1469] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1470] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1471] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1472] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1473] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1474] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1475] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1476] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1477] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1478] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1479] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1480] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1481] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1482] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1483] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1484] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1485] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1486] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1487] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1488] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1489] The following is further disclosed regarding the above embodiment.

[1490] (Claim 1)

[1491] means for acquiring audio data;

[1492] means for converting the acquired voice data into text data;

[1493] means for analyzing the text data to extract dialogue patterns and emotional tones;

[1494] a means for generating feedback based on the analysis results;

[1495] means for presenting the generated feedback to the user;

[1496] A system including:

[1497] (Claim 2)

[1498] 10. The system of claim 1, further comprising means for identifying improvements in open dialogue and active listening based on the analysis results and including the improvements in the feedback.

[1499] (Claim 3)

[1500] 10. The system of claim 1, further comprising means for updating the algorithm based on past analysis results and feedback effects to improve analysis accuracy.

[1501] "Example 1"

[1502] (Claim 1)

[1503] means for collecting audio data;

[1504] A means for converting the collected voice data into text data;

[1505] means for analyzing the text data to extract dialogue patterns and emotional tones;

[1506] a means for generating feedback based on the analysis results;

[1507] means for providing the generated feedback to a user;

[1508] A system including:

[1509] (Claim 2)

[1510] 10. The system of claim 1, further comprising means for evaluating the quality and trends of the dialogue based on the analysis results, identifying areas for improvement in open dialogue and active listening, and including the areas for improvement in feedback.

[1511] (Claim 3)

[1512] 10. The system of claim 1, further comprising means for updating the algorithm based on past analysis results and feedback effects to improve analysis accuracy and feedback quality.

[1513] "Application Example 1"

[1514] (Claim 1)

[1515] means for acquiring audio data;

[1516] means for converting the acquired voice data into text data;

[1517] means for analyzing the text data to extract dialogue patterns and emotional tones;

[1518] a means for generating feedback based on the analysis results;

[1519] means for presenting the generated feedback to the user;

[1520] means for using natural language processing algorithms for generating feedback;

[1521] The means by which these tools are used when training and evaluating staff in brick-and-mortar stores;

[1522] A system including:

[1523] (Claim 2)

[1524] 10. The system of claim 1, further comprising means for identifying areas for improvement in open dialogue and active listening based on the analysis results and including the identified areas for improvement in open dialogue and active listening in the feedback.

[1525] (Claim 3)

[1526] 2. The system of claim 1, further comprising means for updating the algorithm based on past analysis results and feedback effects to improve analysis accuracy.

[1527] "Example 2: Combining Emotion Engines"

[1528] (Claim 1)

[1529] a means for a user to set up the device and launch a recording application prior to a one-on-one session or a coaching session;

[1530] a means for the device to record audio during the session in real time as instructed by a recording application;

[1531] A means for users to upload recordings to a server after the session is over;

[1532] A means for passing the voice data received by the server to a voice recognition engine and converting it into text data;

[1533] A server analyzes the converted text data and extracts dialogue patterns and emotional tones;

[1534] A means for an emotion engine in the server to perform emotion analysis on the analyzed text data and voice data;

[1535] a means for the server to generate feedback based on the analysis result and the emotion recognition result;

[1536] a means for using a generative AI model in generating the feedback; and

[1537] means for presenting the generated feedback to the user;

[1538] A system including:

[1539] (Claim 2)

[1540] 10. The system of claim 1, further comprising means for identifying improvements in open dialogue and active listening based on the analysis results and including the improvements in the feedback.

[1541] (Claim 3)

[1542] 10. The system of claim 1, further comprising means for updating the algorithm based on past analysis results and feedback effects to improve analysis accuracy.

[1543] "Application example 2 when combining emotion engines"

[1544] (Claim 1)

[1545] means for acquiring audio data;

[1546] means for converting the acquired voice data into text data;

[1547] means for analyzing the text data to extract dialogue patterns and emotional tones;

[1548] a means for generating feedback based on the analysis results;

[1549] means for presenting the generated feedback to the user;

[1550] A means of assessing customer service skills based on analytical results and emotional tone;

[1551] A means to provide specific points for improvement in customer service as feedback, and

[1552] A system including:

[1553] (Claim 2)

[1554] 10. The system of claim 1, further comprising means for identifying areas for improvement in open dialogue and active listening based on the analysis results and including such areas in the feedback, as well as means for evaluating a positive tone and positive customer service in customer interactions.

[1555] (Claim 3)

[1556] 10. The system of claim 1, further comprising: means for updating the algorithm based on past analysis results and the effect of feedback to improve analysis accuracy; and means for using a generative AI model to provide feedback and for inputting prompt sentences. [Explanation of symbols]

[1557] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for acquiring audio data; means for converting the acquired voice data into text data; means for analyzing the text data to extract dialogue patterns and emotional tones; a means for generating feedback based on the analysis results; means for presenting the generated feedback to the user; A system including:

2. The system of claim 1 , further comprising means for identifying improvements in open dialogue and active listening based on the analysis results and including the improvements in the feedback.

3. The system of claim 1 further comprising means for updating the algorithm based on past analysis results and feedback effects to improve analysis accuracy.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A