System
A voice-based system analyzes and scores product proposals using AI, providing consistent feedback and training to enhance sales representatives' skills efficiently.
Patent Information
- Application Number
- JP2024141525
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2026-03-06
AI Technical Summary
Traditional role-playing methods for improving sales representatives' product proposal skills are inefficient and lack consistent evaluation, relying heavily on human feedback which can be subjective.
A system that allows users to input product proposals by voice, records and analyzes the voice data using AI to extract keywords, scores the proposals, compares with other users, generates feedback, and provides training modes to improve weak areas.
Enables efficient and consistent evaluation of product proposal skills, allowing new employees to acquire high-level skills quickly through standardized feedback and training.
Smart Images

Figure 2026038190000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Practical role-playing is important in today's sales environment, where there is a demand for improving the skills of sales representatives and agency managers who propose products and for more efficient training of new employees. However, traditional role-playing is often conducted face-to-face, making it difficult to efficiently repeat a large amount of training. Another problem is that the quality of feedback depends on the sales representative, making it difficult to obtain consistent evaluations. The present invention aims to solve these problems. [Means for solving the problem]
[0005] The present invention is a system including a means for a user to input product proposals by voice, a means for recording the voice in real time, a means for transmitting the recorded voice data to a voice recognition AI, a means for analyzing the voice data and extracting important keywords and phrases, a means for scoring the proposal content based on the analysis results, a means for comparing the score with the scores of other users, a means for generating feedback based on the comparison results and presenting it to the user, and a means for providing a training mode to strengthen weak areas. This allows for efficient role-playing and consistent evaluation and feedback, thereby improving product proposal skills and streamlining new employee training.
[0006] Understood. Here is the definition:
[0007] "User" refers to the sales representatives and agency managers who propose products.
[0008] "Product proposal" refers to the act of explaining and proposing the features of a product or service.
[0009] "Voice input means" refers to a function that provides an interface for users to input product proposal details through voice.
[0010] "Means for recording" refers to the function of recording the user's voice as digital data.
[0011] "Voice data" refers to digital data that records the user's voice.
[0012] "Voice recognition AI" refers to an artificial intelligence system that analyzes voice data, converts it into text data, and extracts important information.
[0013] "Means of analysis" refers to the function of extracting important keywords and phrases from audio data and treating them as text data.
[0014] "Means of scoring" refers to the function of evaluating the content of product proposals based on analyzed information and assigning scores.
[0015] "Means for comparison" refers to a function that compares a user's score with the scores of other users and evaluates them relatively.
[0016] The "means for generating feedback" refers to a function that presents areas for improvement and evaluation to the user based on the comparison results and predetermined criteria.
[0017] "Means for presenting" refers to a function for displaying or communicating generated feedback to a user.
[0018] "Means for providing a training mode" refers to a function that identifies areas where the user is weak and provides practice opportunities and instructions to improve those areas.
[0019] These definitions allow each element of the system to be clearly understood. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0022] First, the terms used in the following description will be explained.
[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0028] [First embodiment]
[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0041] This invention is a system in which a user makes a product proposal by voice, the proposal is analyzed using a voice recognition AI, and an evaluation and feedback are provided. A specific embodiment of this system is described below.
[0042] Basic system configuration
[0043] 1. User logs into the app:
[0044] The user accesses the launched application, enters their ID and password on the login screen, and logs in. The server verifies the entered authentication information and permits the login.
[0045] Example: Salesperson A logs in to the app by entering his ID and password.
[0046] 2. Role-playing session begins:
[0047] When the user clicks the "Start Role Play Session" button, the device calls the voice recognition AI and virtual customer module to prepare for the role play session. After preparation is complete, the device prompts the user to start the session.
[0048] Example: When Sales Representative A clicks the "Start Role Play Session" button, the message "Your virtual customer is ready. Please begin" is displayed.
[0049] 3. Role-playing:
[0050] The user makes product proposals to virtual customers by voice. The device records the user's voice and sends it to a voice recognition AI in real time. The server analyzes the voice data and extracts important keywords and phrases.
[0051] Example: When sales representative A says to a virtual customer, "Hello, today I'd like to introduce you to new product X," the audio is recorded and sent to a speech recognition AI.
[0052] 4. Scoring of proposals:
[0053] The server scores the product proposals based on keywords extracted from the voice data. This score is evaluated based on criteria such as the proposal's logic, the amount of information provided, and the degree to which it meets customer needs. Furthermore, the evaluated score is compared with the scores of other users.
[0054] Example: The server extracts keywords such as "New Product X," "Features," and "Benefits," and based on these, assigns an evaluation score of 50 points. The scores of other sales representatives are 60 points, which is relatively low.
[0055] 5. Feedback and training:
[0056] The server evaluates the suggestions and provides specific feedback to the user, including suggestions for improvement. The device displays this feedback to the user. It also provides specific training modes for users to practice repeatedly. This allows users to focus on areas where they are weak, similar to a TOEIC app.
[0057] Example: The server gives feedback that "your proposal is not logical enough," and the device displays this feedback. In addition, the device is presented with the option to start training to improve logical proposal methods.
[0058] System Applications
[0059] This system allows users to effectively improve their product proposal skills. In particular, new employees can acquire high-level skills in a short period of time by repeatedly receiving standardized feedback and training. This will improve the efficiency and results of sales activities.
[0060] As described above, the present invention allows a user to make a product proposal by voice, and the content of the proposal is analyzed and evaluated, thereby improving product proposal skills and making training of new employees more efficient.
[0061] The processing flow will be explained below.
[0062] Step 1:
[0063] The user launches the app, enters their ID and password on the login screen, and clicks the login button.
[0064] Step 2:
[0065] The terminal sends the authentication information (ID and password) entered by the user to the server.
[0066] Step 3:
[0067] The server checks the transmitted authentication information against the database, and if it is correct, generates session information and sends an authentication success message to the terminal.
[0068] Step 4:
[0069] The device receives the authentication success message and displays the home screen.
[0070] Step 5:
[0071] The user clicks the "Start Role Play Session" button on the home screen.
[0072] Step 6:
[0073] The terminal loads the voice recognition AI and virtual customer module, prepares for the role-playing session, and displays a ready message to the user.
[0074] Step 7:
[0075] The user makes product proposals to the virtual customer by voice.
[0076] Step 8:
[0077] The device records the user's voice in real time and sends the voice data to the voice recognition AI.
[0078] Step 9:
[0079] The server uses voice recognition AI to convert the voice data into text data and extract important keywords and phrases.
[0080] Step 10:
[0081] The server then scores the product proposals based on the analyzed text data. Evaluation criteria include logic, information richness, and suitability to customer needs.
[0082] Step 11:
[0083] The server compares the generated score with the scores of other users and calculates the relativity of the rating.
[0084] Step 12:
[0085] The server generates feedback based on the score and evaluation results and sends the feedback to the device.
[0086] Step 13:
[0087] The terminal displays the feedback received from the server to the user, including logical suggestions for improvement and specific instructions.
[0088] Step 14:
[0089] The device identifies the user's weak points and provides a training mode to improve them.
[0090] Step 15:
[0091] The user selects the training mode and again proposes products to the virtual customer.
[0092] Step 16:
[0093] The device will re-record the audio during training and send it to the voice recognition AI.
[0094] Step 17:
[0095] The server analyzes the new audio data, scores it, and generates feedback.
[0096] Step 18:
[0097] The device will display new feedback to the user, who can then choose to train again or start a new role-playing session.
[0098] This system allows users to continuously improve their product proposal skills.
[0099] Example 1
[0100] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0101] Currently, there are insufficient methods for effectively and efficiently improving the product proposal skills required for sales activities. In particular, while there are systems that evaluate and provide feedback on the logic of proposals, the richness of information, and the degree to which they meet customer needs, there are no concrete methods for streamlining the education and training of new sales representatives. This makes it difficult to improve skills in a short period of time, resulting in problems such as not maximizing the efficiency and results of sales activities.
[0102] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0103] In this invention, the server includes: a means for a user to input voice input to propose a product; a means for recording the voice in real time; a means for transmitting the recorded voice data to a voice recognition AI; a means for analyzing the voice data and extracting important keywords and phrases; a means for scoring the proposal based on the analysis results; a means for comparing the score with that of other users; a means for generating feedback based on the comparison results and presenting it to the user; a means for providing a training mode to strengthen weak areas; a means for invoking the voice recognition AI and a virtual customer module and starting a role-playing session; and a means for notifying the user of the evaluation results. This enables effective and efficient improvement of product proposal skills. In particular, it streamlines the education and training of new sales representatives, allowing them to acquire high-level skills in a short period of time.
[0104] "User" refers to the entity that uses the system to propose products and provide education and training.
[0105] "Merchandise" refers to the products or services that are the subject of sale or proposal.
[0106] "Speech recognition AI" refers to a technology or system that analyzes voice data and extracts important keywords and phrases.
[0107] "Virtual Customer Module" refers to software or a system that simulates a virtual customer with whom a user interacts when making a product proposal in a role-play session.
[0108] "Real time" refers to a state in which processing occurs immediately.
[0109] "Scoring" refers to evaluating the content of product proposals and expressing them in numbers or ranks.
[0110] "Feedback" refers to providing specific improvements and advice to the user's product proposal based on the evaluation results.
[0111] "Training Mode" refers to a practice environment designed to help a user improve a particular skill or knowledge.
[0112] A "role-play session" refers to a practice session in which a user simulates making a product proposal to a virtual customer.
[0113] "Evaluation results" refers to the scores and ranks of the scored proposals, as well as the analysis and comments based on them.
[0114] "Important keywords and phrases" refer to the main words and phrases extracted by speech recognition AI from voice data that are necessary for evaluating product proposals.
[0115] This invention is a system in which a user makes a product proposal by voice, the proposal is analyzed using a voice recognition AI, and an evaluation and feedback are provided. The basic configuration of this system is as follows.
[0116] System Configuration
[0117] 1. A user logs into the app
[0118] The user starts the application on the device and enters their ID and password on the login screen. The device sends the user's input to the server, which then verifies the authentication information. If the information matches, the user is permitted to log in and a login success message is displayed on the device.
[0119] Example: A sales representative enters the ID "user123" and password, presses the login button, and is successfully logged in.
[0120] 2. Start of the role-playing session
[0121] When the user clicks the "Start Role-Play Session" button, the device starts the voice recognition AI and virtual customer module. The server notifies the device that the voice recognition AI and virtual customer module are ready, and the device displays instructions to the user to start the session.
[0122] Example: A salesperson presses the "Start Role Play Session" button, and the terminal displays "Your virtual customer is ready. Begin."
[0123] 3. Role-playing
[0124] The user makes product proposals to the virtual customer by voice. The device records the user's voice and sends it to a voice recognition AI in real time. The server analyzes the voice data and extracts important keywords and phrases.
[0125] Example: When a salesperson says, "Hello, today I'd like to introduce you to our new product," the audio is recorded and sent to a voice recognition AI in real time.
[0126] 4. Scoring of proposals
[0127] The server scores the product proposals based on the extracted keywords and phrases. This score is evaluated based on the proposal's logic, the amount of information provided, and the degree to which it meets customer needs. The score is then compared with that of other users, and the results are recorded.
[0128] Example: The server scores the product based on the keywords "new product," "features," and "benefits," and assigns it a score of 50. This is a relatively low score compared to other users' scores (e.g., 60 points).
[0129] 5. Feedback and training
[0130] Based on the evaluation results, the server generates specific feedback and suggestions for improvement for the user. The device displays the generated feedback to the user and provides a training mode for repeated practice. In this training mode, the user can focus on areas where they are weak.
[0131] Example: The server gives feedback that "your proposal lacks logic," and the device displays this information. In addition, the device displays an option to provide training to improve logical proposal methods.
[0132] Hardware and software used
[0133] Device: The computer, smartphone, tablet, etc. that the user uses.
[0134] Server: Cloud server or on-premise server for voice recognition AI and data analysis.
[0135] Speech Recognition AI: An artificial intelligence model that analyzes a user's speech and extracts important keywords and phrases.
[0136] Virtual Customer Module: A software module that acts as a virtual customer.
[0137] Examples of natural language prompts
[0138] "Hello, I'm a sales representative from X Company. Today I'd like to introduce you to our new product. This product has the following features. It also has the following benefits:
[0139] As described above, this invention allows a user to make a product proposal by voice, and the content of the proposal is analyzed and evaluated, thereby improving product proposal skills and making new employee training more efficient.
[0140] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0141] Step 1:
[0142] Input: The user accesses the app's login screen using their ID and password.
[0143] Operation: The user enters their ID and password and presses the login button.
[0144] Processing: The terminal receives the login information and sends it to the server, which checks the received ID and password against the authentication information in its database.
[0145] Output: The server sends an authentication success message to the terminal, and the terminal displays a login success message to the user.
[0146] Step 2:
[0147] Input: User clicks the "Start Role Play Session" button.
[0148] Operation: The device activates the voice recognition AI and virtual customer module and sends a request to the server.
[0149] Processing: The server receives the request, initializes the speech recognition AI and virtual customer module, and prepares the necessary resources.
[0150] Output: The server sends a ready message to the terminal, and the terminal notifies the user: "The virtual customer is ready. Start now."
[0151] Step 3:
[0152] Input: The user initiates a product proposal by voice.
[0153] Action: A user says to a hypothetical customer, "Hello, today I'd like to introduce you to our new product."
[0154] Processing: The device records the user's voice and sends it to the voice recognition AI in real time.
[0155] Output: The audio data is sent to a server, which analyzes it and extracts important keywords and phrases.
[0156] Step 4:
[0157] Input: Voice data analysis results by voice recognition AI.
[0158] How it works: The server receives the analysis results and scores the suggestions based on keywords and phrases.
[0159] Processing: The server generates a score based on the evaluation criteria (logic of the proposal, richness of information, and degree of adaptation to customer needs).
[0160] Output: The generated score is compared with the scores of other users and the results are stored in a database.
[0161] Step 5:
[0162] Input: Scoring and comparison results.
[0163] Action: The server generates feedback based on the comparison results.
[0164] Processing: The server generates specific improvements and advice as feedback and sends it to the device.
[0165] Output: The device displays feedback to the user and provides a training mode for repeated practice.
[0166] Step 6:
[0167] Input: The user initiates the training mode presented.
[0168] Action: The user selects training mode and begins practicing.
[0169] Processing: The device executes the training mode based on the user's selection and sends the user's progress to the server.
[0170] Output: The server tracks the user's progress and again supports training by providing assessment and feedback.
[0171] (Application example 1)
[0172] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0173] In today's brick-and-mortar stores, the ability of salespeople to effectively propose products to customers is extremely important for increasing sales and customer satisfaction. However, it is not easy, especially for new salespeople, to effectively improve their product proposal skills in a short period of time, and there are few standardized training methods. Under these circumstances, a system is needed to efficiently improve salespeople's product proposal skills.
[0174] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0175] In this invention, the server includes means for a user to input product proposals by voice, means for recording the voice in real time, means for transmitting the recorded voice data to a voice recognition AI, means for analyzing the voice data and extracting important keywords and phrases, means for scoring the proposal content based on the analysis results, means for comparing the score with the scores of other users, means for generating feedback based on the comparison results and presenting it to the user, means for providing a training mode for improving weak areas, and means for displaying the feedback to the salesperson on a smartphone. This allows the salesperson to practice product proposals using voice and receive evaluations and feedback in real time.
[0176] "User" refers to an individual who uses the system to train in product proposals.
[0177] "Product proposal" refers to the act of introducing the features and benefits of a product or service to a customer.
[0178] "Voice input means" refers to a device such as a microphone that allows a user to make product suggestions by voice.
[0179] "Recording means" refers to hardware and software for capturing and storing a user's voice.
[0180] "Voice data" refers to the user's recorded voice information.
[0181] "Voice recognition AI" is an artificial intelligence technology that converts voice data into text data and analyzes the content.
[0182] "Means of analysis" refers to using voice recognition AI to extract important keywords and phrases from voice data.
[0183] "Means for scoring proposal content" refers to evaluating the quality of product proposals and converting them into scores.
[0184] "Means for comparing with scores of other users" refers to a function for comparing the scores of proposals made by multiple users.
[0185] "Feedback" refers to information that indicates the evaluation results and areas for improvement regarding the user's product proposal.
[0186] "Training mode" refers to a mode that provides special functionality that allows users to practice their product proposal skills.
[0187] A "smartphone" is a mobile device that has functions such as voice recording, data transmission and reception, and feedback display.
[0188] This invention relates to a smart assistant application that enables sales staff in brick-and-mortar stores to effectively propose products to customers. This system uses voice recognition AI and feedback functions to train sales staff on how to propose products.
[0189] System Configuration
[0190] The basic configuration of the system is as follows:
[0191] 1. User voice input: The user inputs product proposals by voice. A microphone is used for this purpose.
[0192] 2. Recording of voice data: The device (smartphone) records the user's voice in real time and saves it as voice data.
[0193] 3. Sending to speech recognition AI: The recorded voice data is sent from the device to the speech recognition AI, which converts the voice data into text data and analyzes the content.
[0194] 4. Data analysis: The server uses voice recognition AI to analyze the text data and extract important keywords and phrases.
[0195] 5. Scoring of proposal content: The server evaluates and scores the content of the product proposal based on the extracted keywords and phrases.
[0196] 6. Score comparison: The server compares the scores of multiple users and stores the results.
[0197] 7. Feedback generation and presentation: The server generates feedback based on the score comparison results and presents it to the user. This feedback is displayed on the terminal.
[0198] 8. Providing a training mode: The server provides a training mode to help users improve their weak areas. In this mode, specific instructions and examples are provided.
[0199] Program processing
[0200] Hardware:
[0201] Smartphone: Voice recording, data transmission and reception, feedback display
[0202] software:
[0203] Python: Overall program implementation
[0204] speech_recognition library: Audio recording and speech recognition
[0205] some_ai_module: AI module that evaluates product proposal text
[0206] feedback_module: A module that generates feedback based on the evaluation results.
[0207] Specific examples
[0208] For example, when a salesperson proposes new product Y to a customer, the app records the voice and converts it into text using speech recognition AI. The AI module evaluates the quality of the proposal and provides feedback on the proposal skill based on the evaluation.
[0209] Example prompt for a generative AI model:
[0210] Prompt: "You can be more specific about how product Y compares to other products. You can also explain in detail how it meets your customer's needs."
[0211] This invention allows salespeople to practice product proposals using voice and receive evaluations and feedback in real time, which is expected to improve their proposal skills in a short period of time and contribute to increased customer satisfaction.
[0212] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0213] Step 1:
[0214] Users log in to the app and begin voice input of product proposals. As users speak, they describe the features and benefits of the product, and voice data is generated.
[0215] Input: User's voice
[0216] Output: Audio data
[0217] Specific action: The user speaks into the smartphone's microphone.
[0218] Step 2:
[0219] The device records the user's voice in real time and temporarily stores the recorded voice data.
[0220] Input: Audio data
[0221] Output: Recorded audio data
[0222] Specific operation: The recording function in the smartphone is activated and audio data is captured.
[0223] Step 3:
[0224] The device sends the recorded voice data to the voice recognition AI, which converts the voice data into text data.
[0225] Input: Recorded audio data
[0226] Output: Text data
[0227] How it works: The device sends voice data to a cloud-based voice recognition AI service, which analyzes the voice and converts it into text.
[0228] Step 4:
[0229] The server analyzes the text data sent by the voice recognition AI, extracts important keywords and phrases, and stores the results.
[0230] Input: Text data
[0231] Output: Extracted keywords and phrases
[0232] What it does: Speech recognition AI analyzes the text and picks out important keywords related to your business.
[0233] Step 5:
[0234] The server evaluates and scores the product proposals based on the extracted keywords and phrases, and saves the scoring results.
[0235] Input: Extracted keywords or phrases
[0236] Output: Scoring result (evaluation score)
[0237] What it does: An internal server evaluation algorithm scores the usefulness and persuasiveness of keywords.
[0238] Step 6:
[0239] The server compares the scoring results with the scores of other users and stores the results.
[0240] Input: Scoring results
[0241] Output: Comparison result
[0242] What it does: The server compares the current user's score with the scores of other users in its historical database.
[0243] Step 7:
[0244] The server generates feedback based on the score comparison result and sends the feedback to the terminal, where it is displayed.
[0245] Input: Comparison result
[0246] Output: Feedback
[0247] What it does: Based on the comparison results, the feedback generation module will write down in detail what the user needs to improve and what they should praise. The details will be displayed on the smartphone.
[0248] Step 8:
[0249] A training mode is provided to help users improve their weak points through user input, with specific instructions and examples provided for users to practice again and again.
[0250] Input: Feedback
[0251] Output: Improvement and training plans
[0252] What it does: The training mode starts, providing guidance and examples to help users improve their suggestion skills.
[0253] An example of this prompt would be:
[0254] Example prompt sentence:
[0255] Prompt: "You can be more specific about how product Y compares to other products. You can also explain in detail how it meets your customer's needs."
[0256] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0257] This invention is a system in which users make product proposals by voice, the contents of the proposal are analyzed using voice recognition AI, and the user's emotional state is also analyzed using an emotion engine, providing detailed feedback and a training mode to improve the user's product proposal skills.
[0258] Basic system configuration
[0259] 1. User logs into the app:
[0260] The user accesses the launched application, enters their ID and password on the login screen, and logs in. The server verifies the entered authentication information and permits the login.
[0261] Example: Salesperson A logs in to the app by entering his ID and password.
[0262] 2. Role-playing session begins:
[0263] When the user clicks the "Start Role Play Session" button, the terminal calls the voice recognition AI and virtual customer module, prepares for the role play session, and displays a message to the user that the session is ready.
[0264] Example: When Sales Representative A clicks the "Start Role Play Session" button, the message "Your virtual customer is ready. Please begin" is displayed.
[0265] 3. Role-playing:
[0266] The user makes product proposals to virtual customers by voice, and the device records the user's voice and transmits it in real time to the voice recognition AI and emotion engine.
[0267] Example: When salesperson A says to a virtual customer, "Hello, today I'd like to introduce you to our new product X," the audio is recorded and sent to a speech recognition AI and emotion engine.
[0268] 4. Suggestion scoring and sentiment analysis:
[0269] The server scores the product proposals based on keywords and phrases extracted from the voice data, using criteria such as logic, information richness, and suitability to customer needs.
[0270] The emotion engine analyzes the user's voice data to recognize their emotional state, and the results of this emotion analysis are reflected in the feedback.
[0271] Example: The server extracts keywords such as "New Product X," "Features," and "Benefits," and based on these, assigns a rating of 50 points. The emotion engine detects the user's stress and tension.
[0272] 5. Feedback and training:
[0273] The server generates feedback for the user based on the score of the proposal and the results of the sentiment analysis, and sends it to the device. The device then displays the feedback to the user. The feedback includes not only the logic, information richness, and responsiveness of the proposal to customer needs, but also advice based on the user's emotional state.
[0274] Example: Feedback such as "Your proposal lacks logic" is given, followed by emotional advice such as "We saw some tension, so let's try a different approach."
[0275] 6. Training mode available:
[0276] The device identifies the user's weak areas and offers training modes to improve them. If the user is emotionally unstable, specific advice on how to relax is also added.
[0277] Example: A training mode is initiated where the user can practice logical suggestions, and advice such as "take repeated deep breaths" and "speak confidently" is displayed to further relax.
[0278] System Applications
[0279] This system not only improves users' product proposal skills, but also allows them to control their emotions during presentations. This allows for more effective proposals and improved sales results. New employees, in particular, can acquire advanced skills in a short period of time thanks to consistent feedback and detailed training modes.
[0280] As described above, this invention allows users to make product proposals verbally, analyzes the content of the proposal, and provides feedback and training that takes into account their emotional state, thereby improving product proposal skills and making new employee training more efficient.
[0281] The processing flow will be explained below.
[0282] Step 1:
[0283] The user launches the app, enters their ID and password on the login screen, and clicks the login button.
[0284] Step 2:
[0285] The terminal sends the authentication information (ID and password) entered by the user to the server.
[0286] Step 3:
[0287] The server checks the transmitted authentication information against the database, and if it is correct, generates session information and sends an authentication success message to the terminal.
[0288] Step 4:
[0289] The device receives the authentication success message and displays the home screen.
[0290] Step 5:
[0291] The user clicks the "Start Role Play Session" button on the home screen.
[0292] Step 6:
[0293] The terminal loads the voice recognition AI and virtual customer module, prepares for the role-playing session, and displays a ready message to the user.
[0294] Step 7:
[0295] The user makes product proposals to the virtual customer by voice.
[0296] Step 8:
[0297] The device records the user's voice in real time and sends the voice data to a voice recognition AI and emotion engine.
[0298] Step 9:
[0299] The server uses voice recognition AI to convert the voice data into text data and extract important keywords and phrases.
[0300] Step 10:
[0301] The server uses an emotion engine to analyze the user's emotional state from the voice data.
[0302] Step 11:
[0303] The server then scores the product proposals based on the analyzed text data. Evaluation criteria include logic, information richness, and suitability to customer needs.
[0304] Step 12:
[0305] The server compares the generated score with the scores of other users and calculates the relativity of the rating.
[0306] Step 13:
[0307] The server generates feedback based on the score and the results of the sentiment analysis.
[0308] Step 14:
[0309] The server transmits the generated feedback to the terminal.
[0310] Step 15:
[0311] The terminal displays the feedback received from the server to the user, which includes evaluation results regarding logic, richness of information, and suitability to customer needs, as well as advice based on the user's emotional state.
[0312] Step 16:
[0313] The device offers specific training modes based on the user's weaknesses.
[0314] Step 17:
[0315] The user selects the training mode and again proposes products to the virtual customer.
[0316] Step 18:
[0317] The device records the audio during training and sends it to the voice recognition AI and emotion engine.
[0318] Step 19:
[0319] The server analyzes the new audio data and performs scoring and sentiment analysis.
[0320] Step 20:
[0321] The server again generates feedback and sends it to the device.
[0322] Step 21:
[0323] The device will display new feedback to the user, who can then choose to train again or start a new role-playing session.
[0324] This system allows users to continuously improve their product proposal skills.
[0325] Example 2
[0326] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0327] Conventional product proposal training systems focus on analyzing users' proposal content and providing feedback. However, due to a lack of feedback and training based on the user's emotional state, they do not adequately improve presentation skills or emotional control abilities when proposing products. As a result, many users feel nervous and stressed when proposing products, which reduces the effectiveness of their proposals.
[0328] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a user to input a product proposal by voice; a means for recording the voice in real time; a means for transmitting the recorded voice data to a voice recognition AI; a means for analyzing the voice data and extracting important keywords and phrases; a means for scoring the proposal content based on the analysis results; a means for comparing the score with the scores of other users; a means for generating feedback based on the comparison results and presenting it to the user; a means for transmitting the voice data to an emotion engine and analyzing the emotional state; a means for reflecting the emotion analysis results in the feedback; a means for providing feedback including advice taking the emotional state into consideration; and a means for providing a training mode to strengthen weak areas. This makes it possible to improve not only the content of the product proposal but also the user's ability to control their emotions during the presentation.
[0329] "User" refers to a person or entity who uses the system to train in product proposals.
[0330] "Product proposal" refers to the act of explaining and proposing the features and benefits of a product or service.
[0331] "Voice input means" refers to a method by which a user provides information to a system using speech.
[0332] "Real-time recording means" refers to techniques or methods for instantly recording the voices spoken by a user.
[0333] "Voice recognition AI" refers to artificial intelligence technology that analyzes input voice and converts it into text data.
[0334] An "emotion engine" refers to algorithms and technologies that analyze a user's emotional state from input voice data.
[0335] "Keyword and phrase extraction means" refers to technology that selects important words and expressions from recorded audio data.
[0336] "Means for scoring proposal content" refers to a method for evaluating the quality of product proposals based on the analysis results and assigning a score.
[0337] "Feedback" refers to the comments and advice provided to users based on analyzed data and scores.
[0338] "Training mode" refers to a function or state in which the system provides training to improve the user's weak areas.
[0339] "Advice that takes into account the user's emotional state" refers to advice provided based on the results of an analysis of the user's emotions.
[0340] "Means of comparison" refers to a method for comparing a user's proposal score with the scores of other users.
[0341] This invention is a system that analyzes the content of product proposals made by users through voice and the emotional state of the users. The system aims to improve users' product proposal skills by providing detailed feedback and training modes using voice recognition AI and an emotion engine.
[0342] First, the user logs in to the application. The user enters their ID and password, which are then authenticated by the server. This process uses a device such as a smartphone or PC, and uses an authentication system such as OAuth or LDAP on the server side.
[0343] Next, the user clicks the "Start Role Play Session" button, which causes the system to launch a voice recognition AI (e.g., Google® Cloud Speech-to-Text) and virtual customer module. This prepares the role play session, and the device displays to the user, "The virtual customer is ready. Please begin."
[0344] When a user makes a product proposal by voice, the device records this voice and sends it in real time to a voice recognition AI and emotion engine (e.g., IBM Watson (registered trademark) Tone Analyzer). Specifically, when a user says, "Hello, today I'd like to introduce you to new product X," the voice is recorded and analyzed.
[0345] The server scores the proposals based on keywords and phrases extracted from the voice data. Evaluation criteria include logic, information richness, and adaptability to customer needs. The server also uses an emotion engine to analyze the user's emotional state, and these results are reflected in the feedback. For example, keywords such as "New Product X," "Features," and "Benefits" are extracted from the voice recognition AI, and a score of 50 is assigned. At the same time, the emotion engine detects the user's state of tension.
[0346] Based on the results of these analyses, the server generates feedback and sends it to the device. The feedback includes the logic of the proposal, the amount of information provided, the degree to which it meets the customer's needs, and emotional advice. Specifically, messages such as "The proposal lacks logic" and "You seem tense, so relax" are displayed on the device.
[0347] The device then identifies the user's weak points and provides a training mode to strengthen them. The user starts the training mode and performs exercises to improve their skills. For example, the device displays a message saying, "You have started a training mode to practice logical proposals," and includes advice such as "Take deep breaths repeatedly" and "Speak with confidence."
[0348] Examples of prompts include "Please enter your ID and password to log in," "Click the button to start the role-playing session," and "Hello, today I'd like to introduce you to our new product X."
[0349] As described above, this invention enables users to improve not only the content of their product proposals but also their emotional control during presentations. This system provides an effective method for improving user skills and streamlining new employee training.
[0350] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0351] Step 1:
[0352] A user logs in to the app.
[0353] Input: The user enters their ID and password.
[0354] Processing: The device sends this authentication information to the server, which verifies it using an authentication system (e.g., OAuth or LDAP). The user's account information is checked against a database.
[0355] Output: If authentication is successful, the user is logged in.
[0356] Specific operation: The user opens the app on their smartphone and enters their ID and password on the login screen. This is then sent from the device to the server, where it is authenticated and, if successful, the user is allowed to log in.
[0357] Step 2:
[0358] The user clicks the "Start Role Play Session" button.
[0359] Input: User clicks the "Start Role Play Session" button.
[0360] Processing: The device calls the speech recognition AI (e.g., Google Cloud Speech-to-Text) and the virtual customer module to prepare for the session.
[0361] Output: The terminal displays "Your virtual customer is ready. Start now."
[0362] Specific operation: When the user clicks the "Start role-playing session" button, the device launches the voice recognition AI and virtual customer module and displays a message on the screen indicating that it is ready.
[0363] Step 3:
[0364] The user makes product suggestions by voice.
[0365] Input: The user makes a product suggestion by voice.
[0366] Processing: The device records the audio and sends it in real time to a speech recognition AI and emotion engine (e.g., IBM Watson Tone Analyzer). The speech recognition AI converts the audio data into text, and the emotion engine analyzes the emotional state.
[0367] Output: Text converted from audio data and sentiment analysis results.
[0368] What it does: When a user says, "Hello, today I'd like to introduce you to new product X," the voice is recorded and the voice data is used for analysis.
[0369] Step 4:
[0370] Analyze the proposal and emotional state.
[0371] Input: Text data from the speech recognition AI and analysis results from the emotion engine.
[0372] Processing: The server scores the suggestions based on keywords and phrases extracted from the speech data, and the emotion engine identifies the emotional state and retrieves the results.
[0373] Output: Suggestion score and sentiment analysis results.
[0374] Specific operation: The server calculates a score based on data obtained from the voice recognition AI (e.g., "New Product X," "Features," and "Benefits"), and the emotion engine identifies emotions such as tension.
[0375] Step 5:
[0376] Provide feedback.
[0377] Input: Suggestion score and sentiment analysis results.
[0378] Processing: The server generates feedback based on this data and sends it to the device. The feedback includes advice based on logic, information richness, responsiveness to customer needs, and emotional state.
[0379] Output: The feedback message.
[0380] Specific behavior: Feedback such as "Your proposal lacks logic" or "You seem nervous, so please relax" will be displayed on the device.
[0381] Step 6:
[0382] Training mode will be implemented.
[0383] Input: Analysis results identifying weak areas and feedback on emotional state.
[0384] Processing: The device offers a training mode to help users improve their weaknesses and also displays specific advice on their emotional state.
[0385] Output: Start of training mode and the accompanying screen display.
[0386] Specific actions: The device will display specific advice such as "You have started a training mode to practice logical suggestions," "Take repeated deep breaths," and "Speak with confidence."
[0387] (Application example 2)
[0388] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0389] Improving on-site response skills is extremely important in modern security services. However, traditional training methods have struggled to effectively support the improvement of individual response abilities and emotional control. In particular, new security guards lack on-site response experience, posing challenges for their performance in situations that require a quick and appropriate response. This calls for improved on-site response quality and more efficient training methods, and a system that solves this problem is needed.
[0390] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0391] In this invention, the server includes: means for a user to input product proposals by voice; means for recording the voice in real time; means for transmitting the recorded voice data to a voice recognition AI; means for analyzing the voice data and extracting important keywords and phrases; means for scoring the proposals based on the analysis results; means for comparing the score with those of other users; means for generating feedback based on the comparison results and presenting it to the user; means for providing a training mode to strengthen weak areas; means for a guard to input on-site response procedures by voice and analyze the contents of the input; and means for providing a training mode to improve on-site response skills based on the analyzed contents. This allows security guards to effectively train their on-site response capabilities and improve their skills, including emotional control.
[0392] "User" refers to the person who operates the system.
[0393] "Merchandise" refers to the products or services that are the subject of sale or proposal.
[0394] "Speech recognition AI" refers to an artificial intelligence algorithm that converts voice data into text data.
[0395] An "emotion engine" refers to software for analyzing a speaker's emotional state from voice data.
[0396] "Virtual customer module" refers to the part of the computer simulation that responds to user suggestions.
[0397] "Scoring" refers to the process of evaluating proposal content and assigning it a score.
[0398] "Feedback" refers to information that provides an evaluation of a user's statements or actions and advice for improvement.
[0399] "Training mode" refers to a practice mode that a user performs to improve a particular skill.
[0400] A "security guard" is a security guard who works to ensure safety on-site.
[0401] "Scene response procedures" refer to a set of steps that outline what a security guard should do in a particular situation.
[0402] "Analysis" refers to the process of analyzing data in detail and extracting useful information.
[0403] "Skills" refer to the techniques and abilities required to perform a particular task or job.
[0404] A specific system for implementing the present invention is a training system whose main purpose is to help security guards improve their on-site response skills. Details of this system are described below.
[0405] Basic system configuration
[0406] 1. Login function:
[0407] The security guard accesses the launched application and enters the ID and password on the login screen. The server verifies the entered authentication information and allows the login.
[0408] Example: A security guard logs into a system by entering an ID and password.
[0409] 2. Begin the scenario session:
[0410] When the security guard clicks the "Start Scenario Session" button, the terminal calls the voice recognition AI and virtual scene module to prepare for the scenario session, and displays a message to the security guard that the session is ready.
[0411] Example: When a security guard clicks the "Start Scenario Session" button, the message "The virtual scene is ready. Please begin" appears.
[0412] 3. Scenario implementation:
[0413] The security guard will then verbally explain the on-site response procedures to the virtual scene, and the device will record the security guard's voice and transmit it to the voice recognition AI and emotion engine in real time.
[0414] Example: When a security guard says to a virtual scene, "I have spotted a suspicious person. I will leave immediately and call the police," the audio is recorded and sent to a speech recognition AI and emotion engine.
[0415] 4. Suggestion scoring and sentiment analysis:
[0416] The server then scores the response based on keywords and phrases extracted from the voice data, with evaluation criteria including logic, accuracy of information, and appropriateness of the response.
[0417] The emotion engine analyzes the security guard's voice data to recognize their emotional state, and the results of this emotion analysis are also reflected in the feedback.
[0418] Example: The server extracts keywords such as "suspicious person," "police," and "report," and based on that, assigns a rating of 70 points. The emotion engine detects the calmness of the security guard.
[0419] 5. Feedback and training:
[0420] The server generates feedback for the security guard based on the score of the suggestion and the results of the emotion analysis, and sends it to the terminal. The terminal then displays the feedback to the security guard. The feedback includes not only the logic of the response, the accuracy of the information, and the appropriateness of the response, but also advice based on the security guard's emotional state.
[0421] Example: Feedback such as "The logic of your response was appropriate" is given, along with emotional advice such as "You showed some composure, so keep it up."
[0422] 6. Training mode available:
[0423] The device identifies weak areas of the guard and offers a training mode to improve them. If the guard is emotionally unstable, specific advice on how to relax is also added.
[0424] Example: A security guard enters a training mode to practice how to call the police, and offers advice such as "take repeated deep breaths" and "calmly assess the situation" to help the guard relax.
[0425] Hardware and software used
[0426] Hardware: Smartphone, head-mounted display
[0427] Software: SpeechRecognition library, Transformer model
[0428] Prompt Sentence Examples
[0429] What you said: A suspicious person has been spotted. Please leave immediately and call the police.
[0430] Emotional state: POSITIVE
[0431] Expected feedback: Your response was logical and appropriate. You showed some composure, so keep it up.
[0432] The above is the details of the specific embodiment for carrying out the present invention.
[0433] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0434] Step 1:
[0435] Log in
[0436] Input: The user enters the ID and password.
[0437] Operation: The device sends the entered ID and password to the server.
[0438] Data processing and calculation: The server checks the authentication information and compares it with the database.
[0439] Output: If authentication is successful, you will be logged into the system.
[0440] Step 2:
[0441] Starting a Scenario Session
[0442] Input: User clicks the "Start Scenario Session" button.
[0443] Operation: The terminal calls the voice recognition AI and virtual field module to prepare for the scenario session.
[0444] Data processing and calculation: When preparation is complete, the server generates a ready message.
[0445] Output: The terminal displays the message "Virtual site is ready, start now."
[0446] Step 3:
[0447] Scenario implementation
[0448] Input: The user speaks the on-site response procedures into the virtual scene.
[0449] How it works: The device records the user's voice and sends it to the speech recognition AI and emotion engine in real time.
[0450] Data processing and calculation: Voice data is converted into text data, and then an emotion engine performs emotion analysis.
[0451] Output: Text data and emotional state information are generated.
[0452] Step 4:
[0453] Suggestion content scoring and sentiment analysis
[0454] Input: Text data and emotional state information generated in step 3.
[0455] How it works: The server scores the content of on-site responses based on keywords and phrases extracted from the voice data. The emotion engine analyzes the emotional state.
[0456] Data processing and calculation: Logic, accuracy of information, and appropriate response are scored according to the evaluation criteria, and emotional state is also reflected in the scoring.
[0457] Output: Scoring results and sentiment analysis results are generated.
[0458] Step 5:
[0459] Generate feedback
[0460] Input: The scoring and sentiment analysis results generated in step 4.
[0461] How it works: The server generates feedback based on the score of the suggestions and the results of sentiment analysis.
[0462] Data processing and calculation: Create specific evaluations and advice for improvement for users.
[0463] Output: A feedback message is generated and sent to the terminal.
[0464] Step 6:
[0465] View Feedback
[0466] Input: The feedback message generated in step 5.
[0467] Action: The device displays a feedback message.
[0468] Output: The user sees the evaluation results and advice for improvement.
[0469] Step 7:
[0470] Training mode available
[0471] Input: The feedback message the user confirmed in step 6.
[0472] How it works: Identifies the user's weak areas and offers training modes.
[0473] Data processing and calculation: The server generates strengthening training content based on the user's weaknesses. In some cases, it also adds advice on relaxation.
[0474] Output: Training mode is initiated and specific instructions and advice are displayed to the user.
[0475] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0476] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0477] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0478] [Second embodiment]
[0479] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0480] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0481] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0482] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0483] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0484] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0485] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0486] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0487] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0488] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0489] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0490] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0491] This invention is a system in which a user makes a product proposal by voice, the proposal is analyzed using a voice recognition AI, and an evaluation and feedback are provided. A specific embodiment of this system is described below.
[0492] Basic system configuration
[0493] 1. User logs into the app:
[0494] The user accesses the launched application, enters their ID and password on the login screen, and logs in. The server verifies the entered authentication information and permits the login.
[0495] Example: Salesperson A logs in to the app by entering his ID and password.
[0496] 2. Role-playing session begins:
[0497] When the user clicks the "Start Role Play Session" button, the device calls the voice recognition AI and virtual customer module to prepare for the role play session. After preparation is complete, the device prompts the user to start the session.
[0498] Example: When Sales Representative A clicks the "Start Role Play Session" button, the message "Your virtual customer is ready. Please begin" is displayed.
[0499] 3. Role-playing:
[0500] The user makes product proposals to virtual customers by voice. The device records the user's voice and sends it to a voice recognition AI in real time. The server analyzes the voice data and extracts important keywords and phrases.
[0501] Example: When sales representative A says to a virtual customer, "Hello, today I'd like to introduce you to new product X," the audio is recorded and sent to a speech recognition AI.
[0502] 4. Scoring of proposals:
[0503] The server scores the product proposals based on keywords extracted from the voice data. This score is evaluated based on criteria such as the proposal's logic, the amount of information provided, and the degree to which it meets customer needs. Furthermore, the evaluated score is compared with the scores of other users.
[0504] Example: The server extracts keywords such as "New Product X," "Features," and "Benefits," and based on these, assigns an evaluation score of 50 points. The scores of other sales representatives are 60 points, which is relatively low.
[0505] 5. Feedback and training:
[0506] The server evaluates the suggestions and provides specific feedback to the user, including suggestions for improvement. The device displays this feedback to the user. It also provides specific training modes for users to practice repeatedly. This allows users to focus on areas where they are weak, similar to a TOEIC app.
[0507] Example: The server gives feedback that "your proposal is not logical enough," and the device displays this feedback. In addition, the device is presented with the option to start training to improve logical proposal methods.
[0508] System Applications
[0509] This system allows users to effectively improve their product proposal skills. In particular, new employees can acquire high-level skills in a short period of time by repeatedly receiving standardized feedback and training. This will improve the efficiency and results of sales activities.
[0510] As described above, the present invention allows a user to make a product proposal by voice, and the content of the proposal is analyzed and evaluated, thereby improving product proposal skills and making training of new employees more efficient.
[0511] The processing flow will be explained below.
[0512] Step 1:
[0513] The user launches the app, enters their ID and password on the login screen, and clicks the login button.
[0514] Step 2:
[0515] The terminal sends the authentication information (ID and password) entered by the user to the server.
[0516] Step 3:
[0517] The server checks the transmitted authentication information against the database, and if it is correct, generates session information and sends an authentication success message to the terminal.
[0518] Step 4:
[0519] The device receives the authentication success message and displays the home screen.
[0520] Step 5:
[0521] The user clicks the "Start Role Play Session" button on the home screen.
[0522] Step 6:
[0523] The terminal loads the voice recognition AI and virtual customer module, prepares for the role-playing session, and displays a ready message to the user.
[0524] Step 7:
[0525] The user makes product proposals to the virtual customer by voice.
[0526] Step 8:
[0527] The device records the user's voice in real time and sends the voice data to the voice recognition AI.
[0528] Step 9:
[0529] The server uses voice recognition AI to convert the voice data into text data and extract important keywords and phrases.
[0530] Step 10:
[0531] The server then scores the product proposals based on the analyzed text data. Evaluation criteria include logic, information richness, and suitability to customer needs.
[0532] Step 11:
[0533] The server compares the generated score with the scores of other users and calculates the relativity of the rating.
[0534] Step 12:
[0535] The server generates feedback based on the score and evaluation results and sends the feedback to the device.
[0536] Step 13:
[0537] The terminal displays the feedback received from the server to the user, including logical suggestions for improvement and specific instructions.
[0538] Step 14:
[0539] The device identifies the user's weak points and provides a training mode to improve them.
[0540] Step 15:
[0541] The user selects the training mode and again proposes products to the virtual customer.
[0542] Step 16:
[0543] The device will re-record the audio during training and send it to the voice recognition AI.
[0544] Step 17:
[0545] The server analyzes the new audio data, scores it, and generates feedback.
[0546] Step 18:
[0547] The device will display new feedback to the user, who can then choose to train again or start a new role-playing session.
[0548] This system allows users to continuously improve their product proposal skills.
[0549] Example 1
[0550] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0551] Currently, there are insufficient methods for effectively and efficiently improving the product proposal skills required for sales activities. In particular, while there are systems that evaluate and provide feedback on the logic of proposals, the richness of information, and the degree to which they meet customer needs, there are no concrete methods for streamlining the education and training of new sales representatives. This makes it difficult to improve skills in a short period of time, resulting in problems such as not maximizing the efficiency and results of sales activities.
[0552] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0553] In this invention, the server includes: a means for a user to input voice input to propose a product; a means for recording the voice in real time; a means for transmitting the recorded voice data to a voice recognition AI; a means for analyzing the voice data and extracting important keywords and phrases; a means for scoring the proposal based on the analysis results; a means for comparing the score with that of other users; a means for generating feedback based on the comparison results and presenting it to the user; a means for providing a training mode to strengthen weak areas; a means for invoking the voice recognition AI and a virtual customer module and starting a role-playing session; and a means for notifying the user of the evaluation results. This enables effective and efficient improvement of product proposal skills. In particular, it streamlines the education and training of new sales representatives, allowing them to acquire high-level skills in a short period of time.
[0554] "User" refers to the entity that uses the system to propose products and provide education and training.
[0555] "Merchandise" refers to the products or services that are the subject of sale or proposal.
[0556] "Speech recognition AI" refers to a technology or system that analyzes voice data and extracts important keywords and phrases.
[0557] "Virtual Customer Module" refers to software or a system that simulates a virtual customer with whom a user interacts when making a product proposal in a role-play session.
[0558] "Real time" refers to a state in which processing occurs immediately.
[0559] "Scoring" refers to evaluating the content of product proposals and expressing them in numbers or ranks.
[0560] "Feedback" refers to providing specific improvements and advice to the user's product proposal based on the evaluation results.
[0561] "Training Mode" refers to a practice environment designed to help a user improve a particular skill or knowledge.
[0562] A "role-play session" refers to a practice session in which a user simulates making a product proposal to a virtual customer.
[0563] "Evaluation results" refers to the scores and ranks of the scored proposals, as well as the analysis and comments based on them.
[0564] "Important keywords and phrases" refer to the main words and phrases extracted by speech recognition AI from voice data that are necessary for evaluating product proposals.
[0565] This invention is a system in which a user makes a product proposal by voice, the proposal is analyzed using a voice recognition AI, and an evaluation and feedback are provided. The basic configuration of this system is as follows.
[0566] System Configuration
[0567] 1. A user logs into the app
[0568] The user starts the application on the device and enters their ID and password on the login screen. The device sends the user's input to the server, which then verifies the authentication information. If the information matches, the user is permitted to log in and a login success message is displayed on the device.
[0569] Example: A sales representative enters the ID "user123" and password, presses the login button, and is successfully logged in.
[0570] 2. Start of the role-playing session
[0571] When the user clicks the "Start Role-Play Session" button, the device starts the voice recognition AI and virtual customer module. The server notifies the device that the voice recognition AI and virtual customer module are ready, and the device displays instructions to the user to start the session.
[0572] Example: A salesperson presses the "Start Role Play Session" button, and the terminal displays "Your virtual customer is ready. Begin."
[0573] 3. Role-playing
[0574] The user makes product proposals to the virtual customer by voice. The device records the user's voice and sends it to a voice recognition AI in real time. The server analyzes the voice data and extracts important keywords and phrases.
[0575] Example: When a salesperson says, "Hello, today I'd like to introduce you to our new product," the audio is recorded and sent to a voice recognition AI in real time.
[0576] 4. Scoring of proposals
[0577] The server scores the product proposals based on the extracted keywords and phrases. This score is evaluated based on the proposal's logic, the amount of information provided, and the degree to which it meets customer needs. The score is then compared with that of other users, and the results are recorded.
[0578] Example: The server scores the product based on the keywords "new product," "features," and "benefits," and assigns it a score of 50. This is a relatively low score compared to other users' scores (e.g., 60 points).
[0579] 5. Feedback and training
[0580] Based on the evaluation results, the server generates specific feedback and suggestions for improvement for the user. The device displays the generated feedback to the user and provides a training mode for repeated practice. In this training mode, the user can focus on areas where they are weak.
[0581] Example: The server gives feedback that "your proposal lacks logic," and the device displays this information. In addition, the device displays an option to provide training to improve logical proposal methods.
[0582] Hardware and software used
[0583] Device: The computer, smartphone, tablet, etc. that the user uses.
[0584] Server: Cloud server or on-premise server for voice recognition AI and data analysis.
[0585] Speech Recognition AI: An artificial intelligence model that analyzes a user's speech and extracts important keywords and phrases.
[0586] Virtual Customer Module: A software module that acts as a virtual customer.
[0587] Examples of natural language prompts
[0588] "Hello, I'm a sales representative from X Company. Today I'd like to introduce you to our new product. This product has the following features. It also has the following benefits:
[0589] As described above, this invention allows a user to make a product proposal by voice, and the content of the proposal is analyzed and evaluated, thereby improving product proposal skills and making new employee training more efficient.
[0590] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0591] Step 1:
[0592] Input: The user accesses the app's login screen using their ID and password.
[0593] Operation: The user enters their ID and password and presses the login button.
[0594] Processing: The terminal receives the login information and sends it to the server, which checks the received ID and password against the authentication information in its database.
[0595] Output: The server sends an authentication success message to the terminal, and the terminal displays a login success message to the user.
[0596] Step 2:
[0597] Input: User clicks the "Start Role Play Session" button.
[0598] Operation: The device activates the voice recognition AI and virtual customer module and sends a request to the server.
[0599] Processing: The server receives the request, initializes the speech recognition AI and virtual customer module, and prepares the necessary resources.
[0600] Output: The server sends a ready message to the terminal, and the terminal notifies the user: "The virtual customer is ready. Start now."
[0601] Step 3:
[0602] Input: The user initiates a product proposal by voice.
[0603] Action: A user says to a hypothetical customer, "Hello, today I'd like to introduce you to our new product."
[0604] Processing: The device records the user's voice and sends it to the voice recognition AI in real time.
[0605] Output: The audio data is sent to a server, which analyzes it and extracts important keywords and phrases.
[0606] Step 4:
[0607] Input: Voice data analysis results by voice recognition AI.
[0608] How it works: The server receives the analysis results and scores the suggestions based on keywords and phrases.
[0609] Processing: The server generates a score based on the evaluation criteria (logic of the proposal, richness of information, and degree of adaptation to customer needs).
[0610] Output: The generated score is compared with the scores of other users and the results are stored in a database.
[0611] Step 5:
[0612] Input: Scoring and comparison results.
[0613] Action: The server generates feedback based on the comparison results.
[0614] Processing: The server generates specific improvements and advice as feedback and sends it to the device.
[0615] Output: The device displays feedback to the user and provides a training mode for repeated practice.
[0616] Step 6:
[0617] Input: The user initiates the training mode presented.
[0618] Action: The user selects training mode and begins practicing.
[0619] Processing: The device executes the training mode based on the user's selection and sends the user's progress to the server.
[0620] Output: The server tracks the user's progress and again supports training by providing assessment and feedback.
[0621] (Application example 1)
[0622] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0623] In today's brick-and-mortar stores, the ability of salespeople to effectively propose products to customers is extremely important for increasing sales and customer satisfaction. However, it is not easy, especially for new salespeople, to effectively improve their product proposal skills in a short period of time, and there are few standardized training methods. Under these circumstances, a system is needed to efficiently improve salespeople's product proposal skills.
[0624] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0625] In this invention, the server includes means for a user to input product proposals by voice, means for recording the voice in real time, means for transmitting the recorded voice data to a voice recognition AI, means for analyzing the voice data and extracting important keywords and phrases, means for scoring the proposal content based on the analysis results, means for comparing the score with the scores of other users, means for generating feedback based on the comparison results and presenting it to the user, means for providing a training mode for improving weak areas, and means for displaying the feedback to the salesperson on a smartphone. This allows the salesperson to practice product proposals using voice and receive evaluations and feedback in real time.
[0626] "User" refers to an individual who uses the system to train in product proposals.
[0627] "Product proposal" refers to the act of introducing the features and benefits of a product or service to a customer.
[0628] "Voice input means" refers to a device such as a microphone that allows a user to make product suggestions by voice.
[0629] "Recording means" refers to hardware and software for capturing and storing a user's voice.
[0630] "Voice data" refers to the user's recorded voice information.
[0631] "Voice recognition AI" is an artificial intelligence technology that converts voice data into text data and analyzes the content.
[0632] "Means of analysis" refers to using voice recognition AI to extract important keywords and phrases from voice data.
[0633] "Means for scoring proposal content" refers to evaluating the quality of product proposals and converting them into scores.
[0634] "Means for comparing with scores of other users" refers to a function for comparing the scores of proposals made by multiple users.
[0635] "Feedback" refers to information that indicates the evaluation results and areas for improvement regarding the user's product proposal.
[0636] "Training mode" refers to a mode that provides special functionality that allows users to practice their product proposal skills.
[0637] A "smartphone" is a mobile device that has functions such as voice recording, data transmission and reception, and feedback display.
[0638] This invention relates to a smart assistant application that enables sales staff in brick-and-mortar stores to effectively propose products to customers. This system uses voice recognition AI and feedback functions to train sales staff on how to propose products.
[0639] System Configuration
[0640] The basic configuration of the system is as follows:
[0641] 1. User voice input: The user inputs product proposals by voice. A microphone is used for this purpose.
[0642] 2. Recording of voice data: The device (smartphone) records the user's voice in real time and saves it as voice data.
[0643] 3. Sending to speech recognition AI: The recorded voice data is sent from the device to the speech recognition AI, which converts the voice data into text data and analyzes the content.
[0644] 4. Data analysis: The server uses voice recognition AI to analyze the text data and extract important keywords and phrases.
[0645] 5. Scoring of proposal content: The server evaluates and scores the content of the product proposal based on the extracted keywords and phrases.
[0646] 6. Score comparison: The server compares the scores of multiple users and stores the results.
[0647] 7. Feedback generation and presentation: The server generates feedback based on the score comparison results and presents it to the user. This feedback is displayed on the terminal.
[0648] 8. Providing a training mode: The server provides a training mode to help users improve their weak areas. In this mode, specific instructions and examples are provided.
[0649] Program processing
[0650] Hardware:
[0651] Smartphone: Voice recording, data transmission and reception, feedback display
[0652] software:
[0653] Python: Overall program implementation
[0654] speech_recognition library: Audio recording and speech recognition
[0655] some_ai_module: AI module that evaluates product proposal text
[0656] feedback_module: A module that generates feedback based on the evaluation results.
[0657] Specific examples
[0658] For example, when a salesperson proposes new product Y to a customer, the app records the voice and converts it into text using speech recognition AI. The AI module evaluates the quality of the proposal and provides feedback on the proposal skill based on the evaluation.
[0659] Example prompt for a generative AI model:
[0660] Prompt: "You can be more specific about how product Y compares to other products. You can also explain in detail how it meets your customer's needs."
[0661] This invention allows salespeople to practice product proposals using voice and receive evaluations and feedback in real time, which is expected to improve their proposal skills in a short period of time and contribute to increased customer satisfaction.
[0662] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0663] Step 1:
[0664] Users log in to the app and begin voice input of product proposals. As users speak, they describe the features and benefits of the product, and voice data is generated.
[0665] Input: User's voice
[0666] Output: Audio data
[0667] Specific action: The user speaks into the smartphone's microphone.
[0668] Step 2:
[0669] The device records the user's voice in real time and temporarily stores the recorded voice data.
[0670] Input: Audio data
[0671] Output: Recorded audio data
[0672] Specific operation: The recording function in the smartphone is activated and audio data is captured.
[0673] Step 3:
[0674] The device sends the recorded voice data to the voice recognition AI, which converts the voice data into text data.
[0675] Input: Recorded audio data
[0676] Output: Text data
[0677] How it works: The device sends voice data to a cloud-based voice recognition AI service, which analyzes the voice and converts it into text.
[0678] Step 4:
[0679] The server analyzes the text data sent by the voice recognition AI, extracts important keywords and phrases, and stores the results.
[0680] Input: Text data
[0681] Output: Extracted keywords and phrases
[0682] What it does: Speech recognition AI analyzes the text and picks out important keywords related to your business.
[0683] Step 5:
[0684] The server evaluates and scores the product proposals based on the extracted keywords and phrases, and saves the scoring results.
[0685] Input: Extracted keywords or phrases
[0686] Output: Scoring result (evaluation score)
[0687] What it does: An internal server evaluation algorithm scores the usefulness and persuasiveness of keywords.
[0688] Step 6:
[0689] The server compares the scoring results with the scores of other users and stores the results.
[0690] Input: Scoring results
[0691] Output: Comparison result
[0692] What it does: The server compares the current user's score with the scores of other users in its historical database.
[0693] Step 7:
[0694] The server generates feedback based on the score comparison result and sends the feedback to the terminal, where it is displayed.
[0695] Input: Comparison result
[0696] Output: Feedback
[0697] What it does: Based on the comparison results, the feedback generation module will write down in detail what the user needs to improve and what they should praise. The details will be displayed on the smartphone.
[0698] Step 8:
[0699] A training mode is provided to help users improve their weak points through user input, with specific instructions and examples provided for users to practice again and again.
[0700] Input: Feedback
[0701] Output: Improvement and training plans
[0702] What it does: The training mode starts, providing guidance and examples to help users improve their suggestion skills.
[0703] An example of this prompt would be:
[0704] Example prompt sentence:
[0705] Prompt: "You can be more specific about how product Y compares to other products. You can also explain in detail how it meets your customer's needs."
[0706] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0707] This invention is a system in which users make product proposals by voice, the contents of the proposal are analyzed using voice recognition AI, and the user's emotional state is also analyzed using an emotion engine, providing detailed feedback and a training mode to improve the user's product proposal skills.
[0708] Basic system configuration
[0709] 1. User logs into the app:
[0710] The user accesses the launched application, enters their ID and password on the login screen, and logs in. The server verifies the entered authentication information and permits the login.
[0711] Example: Salesperson A logs in to the app by entering his ID and password.
[0712] 2. Role-playing session begins:
[0713] When the user clicks the "Start Role Play Session" button, the terminal calls the voice recognition AI and virtual customer module, prepares for the role play session, and displays a message to the user that the session is ready.
[0714] Example: When Sales Representative A clicks the "Start Role Play Session" button, the message "Your virtual customer is ready. Please begin" is displayed.
[0715] 3. Role-playing:
[0716] The user makes product proposals to virtual customers by voice, and the device records the user's voice and transmits it in real time to the voice recognition AI and emotion engine.
[0717] Example: When salesperson A says to a virtual customer, "Hello, today I'd like to introduce you to our new product X," the audio is recorded and sent to a speech recognition AI and emotion engine.
[0718] 4. Suggestion scoring and sentiment analysis:
[0719] The server scores the product proposals based on keywords and phrases extracted from the voice data, using criteria such as logic, information richness, and suitability to customer needs.
[0720] The emotion engine analyzes the user's voice data to recognize their emotional state, and the results of this emotion analysis are reflected in the feedback.
[0721] Example: The server extracts keywords such as "New Product X," "Features," and "Benefits," and based on these, assigns a rating of 50 points. The emotion engine detects the user's stress and tension.
[0722] 5. Feedback and training:
[0723] The server generates feedback for the user based on the score of the proposal and the results of the sentiment analysis, and sends it to the device. The device then displays the feedback to the user. The feedback includes not only the logic, information richness, and responsiveness of the proposal to customer needs, but also advice based on the user's emotional state.
[0724] Example: Feedback such as "Your proposal lacks logic" is given, followed by emotional advice such as "We saw some tension, so let's try a different approach."
[0725] 6. Training mode available:
[0726] The device identifies the user's weak areas and offers training modes to improve them. If the user is emotionally unstable, specific advice on how to relax is also added.
[0727] Example: A training mode is initiated where the user can practice logical suggestions, and advice such as "take repeated deep breaths" and "speak confidently" is displayed to further relax.
[0728] System Applications
[0729] This system not only improves users' product proposal skills, but also allows them to control their emotions during presentations. This allows for more effective proposals and improved sales results. New employees, in particular, can acquire advanced skills in a short period of time thanks to consistent feedback and detailed training modes.
[0730] As described above, this invention allows users to make product proposals verbally, analyzes the content of the proposal, and provides feedback and training that takes into account their emotional state, thereby improving product proposal skills and making new employee training more efficient.
[0731] The processing flow will be explained below.
[0732] Step 1:
[0733] The user launches the app, enters their ID and password on the login screen, and clicks the login button.
[0734] Step 2:
[0735] The terminal sends the authentication information (ID and password) entered by the user to the server.
[0736] Step 3:
[0737] The server checks the transmitted authentication information against the database, and if it is correct, generates session information and sends an authentication success message to the terminal.
[0738] Step 4:
[0739] The device receives the authentication success message and displays the home screen.
[0740] Step 5:
[0741] The user clicks the "Start Role Play Session" button on the home screen.
[0742] Step 6:
[0743] The terminal loads the voice recognition AI and virtual customer module, prepares for the role-playing session, and displays a ready message to the user.
[0744] Step 7:
[0745] The user makes product proposals to the virtual customer by voice.
[0746] Step 8:
[0747] The device records the user's voice in real time and sends the voice data to a voice recognition AI and emotion engine.
[0748] Step 9:
[0749] The server uses voice recognition AI to convert the voice data into text data and extract important keywords and phrases.
[0750] Step 10:
[0751] The server uses an emotion engine to analyze the user's emotional state from the voice data.
[0752] Step 11:
[0753] The server then scores the product proposals based on the analyzed text data. Evaluation criteria include logic, information richness, and suitability to customer needs.
[0754] Step 12:
[0755] The server compares the generated score with the scores of other users and calculates the relativity of the rating.
[0756] Step 13:
[0757] The server generates feedback based on the score and the results of the sentiment analysis.
[0758] Step 14:
[0759] The server transmits the generated feedback to the terminal.
[0760] Step 15:
[0761] The terminal displays the feedback received from the server to the user, which includes evaluation results regarding logic, richness of information, and suitability to customer needs, as well as advice based on the user's emotional state.
[0762] Step 16:
[0763] The device offers specific training modes based on the user's weaknesses.
[0764] Step 17:
[0765] The user selects the training mode and again proposes products to the virtual customer.
[0766] Step 18:
[0767] The device records the audio during training and sends it to the voice recognition AI and emotion engine.
[0768] Step 19:
[0769] The server analyzes the new audio data and performs scoring and sentiment analysis.
[0770] Step 20:
[0771] The server again generates feedback and sends it to the device.
[0772] Step 21:
[0773] The device will display new feedback to the user, who can then choose to train again or start a new role-playing session.
[0774] This system allows users to continuously improve their product proposal skills.
[0775] Example 2
[0776] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0777] Conventional product proposal training systems focus on analyzing users' proposal content and providing feedback. However, due to a lack of feedback and training based on the user's emotional state, they do not adequately improve presentation skills or emotional control abilities when proposing products. As a result, many users feel nervous and stressed when proposing products, which reduces the effectiveness of their proposals.
[0778] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a user to input a product proposal by voice; a means for recording the voice in real time; a means for transmitting the recorded voice data to a voice recognition AI; a means for analyzing the voice data and extracting important keywords and phrases; a means for scoring the proposal content based on the analysis results; a means for comparing the score with the scores of other users; a means for generating feedback based on the comparison results and presenting it to the user; a means for transmitting the voice data to an emotion engine and analyzing the emotional state; a means for reflecting the emotion analysis results in the feedback; a means for providing feedback including advice taking the emotional state into consideration; and a means for providing a training mode to strengthen weak areas. This makes it possible to improve not only the content of the product proposal but also the user's ability to control their emotions during the presentation.
[0779] "User" refers to a person or entity who uses the system to train in product proposals.
[0780] "Product proposal" refers to the act of explaining and proposing the features and benefits of a product or service.
[0781] "Voice input means" refers to a method by which a user provides information to a system using speech.
[0782] "Real-time recording means" refers to techniques or methods for instantly recording the voices spoken by a user.
[0783] "Voice recognition AI" refers to artificial intelligence technology that analyzes input voice and converts it into text data.
[0784] An "emotion engine" refers to algorithms and technologies that analyze a user's emotional state from input voice data.
[0785] "Keyword and phrase extraction means" refers to technology that selects important words and expressions from recorded audio data.
[0786] "Means for scoring proposal content" refers to a method for evaluating the quality of product proposals based on the analysis results and assigning a score.
[0787] "Feedback" refers to the comments and advice provided to users based on analyzed data and scores.
[0788] "Training mode" refers to a function or state in which the system provides training to improve the user's weak areas.
[0789] "Advice that takes into account the user's emotional state" refers to advice provided based on the results of an analysis of the user's emotions.
[0790] "Means of comparison" refers to a method for comparing a user's proposal score with the scores of other users.
[0791] This invention is a system that analyzes the content of product proposals made by users through voice and the emotional state of the users. The system aims to improve users' product proposal skills by providing detailed feedback and training modes using voice recognition AI and an emotion engine.
[0792] First, the user logs in to the application. The user enters their ID and password, which are then authenticated by the server. This process uses a device such as a smartphone or PC, and uses an authentication system such as OAuth or LDAP on the server side.
[0793] Next, the user clicks the "Start Role Play Session" button, which causes the system to launch the voice recognition AI (e.g., Google Cloud Speech-to-Text) and virtual customer module. This prepares the role play session, and the device displays to the user, "The virtual customer is ready. Please begin."
[0794] When a user makes a product proposal by voice, the device records this voice and sends it in real time to a voice recognition AI and emotion engine (e.g., IBM Watson Tone Analyzer). Specifically, when a user says, "Hello, today I'd like to introduce you to new product X," the voice is recorded and analyzed.
[0795] The server scores the proposals based on keywords and phrases extracted from the voice data. Evaluation criteria include logic, information richness, and adaptability to customer needs. The server also uses an emotion engine to analyze the user's emotional state, and these results are reflected in the feedback. For example, keywords such as "New Product X," "Features," and "Benefits" are extracted from the voice recognition AI, and a score of 50 is assigned. At the same time, the emotion engine detects the user's state of tension.
[0796] Based on the results of these analyses, the server generates feedback and sends it to the device. The feedback includes the logic of the proposal, the amount of information provided, the degree to which it meets the customer's needs, and emotional advice. Specifically, messages such as "The proposal lacks logic" and "You seem tense, so relax" are displayed on the device.
[0797] The device then identifies the user's weak points and provides a training mode to strengthen them. The user starts the training mode and performs exercises to improve their skills. For example, the device displays a message saying, "You have started a training mode to practice logical proposals," and includes advice such as "Take deep breaths repeatedly" and "Speak with confidence."
[0798] Examples of prompts include "Please enter your ID and password to log in," "Click the button to start the role-playing session," and "Hello, today I'd like to introduce you to our new product X."
[0799] As described above, this invention enables users to improve not only the content of their product proposals but also their emotional control during presentations. This system provides an effective method for improving user skills and streamlining new employee training.
[0800] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0801] Step 1:
[0802] A user logs in to the app.
[0803] Input: The user enters their ID and password.
[0804] Processing: The device sends this authentication information to the server, which verifies it using an authentication system (e.g., OAuth or LDAP). The user's account information is checked against a database.
[0805] Output: If authentication is successful, the user is logged in.
[0806] Specific operation: The user opens the app on their smartphone and enters their ID and password on the login screen. This is then sent from the device to the server, where it is authenticated and, if successful, the user is allowed to log in.
[0807] Step 2:
[0808] The user clicks the "Start Role Play Session" button.
[0809] Input: User clicks the "Start Role Play Session" button.
[0810] Processing: The device calls the speech recognition AI (e.g., Google Cloud Speech-to-Text) and the virtual customer module to prepare for the session.
[0811] Output: The terminal displays "Your virtual customer is ready. Start now."
[0812] Specific operation: When the user clicks the "Start role-playing session" button, the device launches the voice recognition AI and virtual customer module and displays a message on the screen indicating that it is ready.
[0813] Step 3:
[0814] The user makes product suggestions by voice.
[0815] Input: The user makes a product suggestion by voice.
[0816] Processing: The device records the audio and sends it in real time to a speech recognition AI and emotion engine (e.g., IBM Watson Tone Analyzer). The speech recognition AI converts the audio data into text, and the emotion engine analyzes the emotional state.
[0817] Output: Text converted from audio data and sentiment analysis results.
[0818] What it does: When a user says, "Hello, today I'd like to introduce you to new product X," the voice is recorded and the voice data is used for analysis.
[0819] Step 4:
[0820] Analyze the proposal and emotional state.
[0821] Input: Text data from the speech recognition AI and analysis results from the emotion engine.
[0822] Processing: The server scores the suggestions based on keywords and phrases extracted from the speech data, and the emotion engine identifies the emotional state and retrieves the results.
[0823] Output: Suggestion score and sentiment analysis results.
[0824] Specific operation: The server calculates a score based on data obtained from the voice recognition AI (e.g., "New Product X," "Features," and "Benefits"), and the emotion engine identifies emotions such as tension.
[0825] Step 5:
[0826] Provide feedback.
[0827] Input: Suggestion score and sentiment analysis results.
[0828] Processing: The server generates feedback based on this data and sends it to the device. The feedback includes advice based on logic, information richness, responsiveness to customer needs, and emotional state.
[0829] Output: The feedback message.
[0830] Specific behavior: Feedback such as "Your proposal lacks logic" or "You seem nervous, so please relax" will be displayed on the device.
[0831] Step 6:
[0832] Training mode will be implemented.
[0833] Input: Analysis results identifying weak areas and feedback on emotional state.
[0834] Processing: The device offers a training mode to help users improve their weaknesses and also displays specific advice on their emotional state.
[0835] Output: Start of training mode and the accompanying screen display.
[0836] Specific actions: The device will display specific advice such as "You have started a training mode to practice logical suggestions," "Take repeated deep breaths," and "Speak with confidence."
[0837] (Application example 2)
[0838] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0839] Improving on-site response skills is extremely important in modern security services. However, traditional training methods have struggled to effectively support the improvement of individual response abilities and emotional control. In particular, new security guards lack on-site response experience, posing challenges for their performance in situations that require a quick and appropriate response. This calls for improved on-site response quality and more efficient training methods, and a system that solves this problem is needed.
[0840] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0841] In this invention, the server includes: means for a user to input product proposals by voice; means for recording the voice in real time; means for transmitting the recorded voice data to a voice recognition AI; means for analyzing the voice data and extracting important keywords and phrases; means for scoring the proposals based on the analysis results; means for comparing the score with those of other users; means for generating feedback based on the comparison results and presenting it to the user; means for providing a training mode to strengthen weak areas; means for a guard to input on-site response procedures by voice and analyze the contents of the input; and means for providing a training mode to improve on-site response skills based on the analyzed contents. This allows security guards to effectively train their on-site response capabilities and improve their skills, including emotional control.
[0842] "User" refers to the person who operates the system.
[0843] "Merchandise" refers to the products or services that are the subject of sale or proposal.
[0844] "Speech recognition AI" refers to an artificial intelligence algorithm that converts voice data into text data.
[0845] An "emotion engine" refers to software for analyzing a speaker's emotional state from voice data.
[0846] "Virtual customer module" refers to the part of the computer simulation that responds to user suggestions.
[0847] "Scoring" refers to the process of evaluating proposal content and assigning it a score.
[0848] "Feedback" refers to information that provides an evaluation of a user's statements or actions and advice for improvement.
[0849] "Training mode" refers to a practice mode that a user performs to improve a particular skill.
[0850] A "security guard" is a security guard who works to ensure safety on-site.
[0851] "Scene response procedures" refer to a set of steps that outline what a security guard should do in a particular situation.
[0852] "Analysis" refers to the process of analyzing data in detail and extracting useful information.
[0853] "Skills" refer to the techniques and abilities required to perform a particular task or job.
[0854] A specific system for implementing the present invention is a training system whose main purpose is to help security guards improve their on-site response skills. Details of this system are described below.
[0855] Basic system configuration
[0856] 1. Login function:
[0857] The security guard accesses the launched application and enters the ID and password on the login screen. The server verifies the entered authentication information and allows the login.
[0858] Example: A security guard logs into a system by entering an ID and password.
[0859] 2. Begin the scenario session:
[0860] When the security guard clicks the "Start Scenario Session" button, the terminal calls the voice recognition AI and virtual scene module to prepare for the scenario session, and displays a message to the security guard that the session is ready.
[0861] Example: When a security guard clicks the "Start Scenario Session" button, the message "The virtual scene is ready. Please begin" appears.
[0862] 3. Scenario implementation:
[0863] The security guard will then verbally explain the on-site response procedures to the virtual scene, and the device will record the security guard's voice and transmit it to the voice recognition AI and emotion engine in real time.
[0864] Example: When a security guard says to a virtual scene, "I have spotted a suspicious person. I will leave immediately and call the police," the audio is recorded and sent to a speech recognition AI and emotion engine.
[0865] 4. Suggestion scoring and sentiment analysis:
[0866] The server then scores the response based on keywords and phrases extracted from the voice data, with evaluation criteria including logic, accuracy of information, and appropriateness of the response.
[0867] The emotion engine analyzes the security guard's voice data to recognize their emotional state, and the results of this emotion analysis are also reflected in the feedback.
[0868] Example: The server extracts keywords such as "suspicious person," "police," and "report," and based on that, assigns a rating of 70 points. The emotion engine detects the calmness of the security guard.
[0869] 5. Feedback and training:
[0870] The server generates feedback for the security guard based on the score of the suggestion and the results of the emotion analysis, and sends it to the terminal. The terminal then displays the feedback to the security guard. The feedback includes not only the logic of the response, the accuracy of the information, and the appropriateness of the response, but also advice based on the security guard's emotional state.
[0871] Example: Feedback such as "The logic of your response was appropriate" is given, along with emotional advice such as "You showed some composure, so keep it up."
[0872] 6. Training mode available:
[0873] The device identifies weak areas of the guard and offers a training mode to improve them. If the guard is emotionally unstable, specific advice on how to relax is also added.
[0874] Example: A security guard enters a training mode to practice how to call the police, and offers advice such as "take repeated deep breaths" and "calmly assess the situation" to help the guard relax.
[0875] Hardware and software used
[0876] Hardware: Smartphone, head-mounted display
[0877] Software: SpeechRecognition library, Transformer model
[0878] Prompt Sentence Examples
[0879] What you said: A suspicious person has been spotted. Please leave immediately and call the police.
[0880] Emotional state: POSITIVE
[0881] Expected feedback: Your response was logical and appropriate. You showed some composure, so keep it up.
[0882] The above is the details of the specific embodiment for carrying out the present invention.
[0883] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0884] Step 1:
[0885] Log in
[0886] Input: The user enters the ID and password.
[0887] Operation: The device sends the entered ID and password to the server.
[0888] Data processing and calculation: The server checks the authentication information and compares it with the database.
[0889] Output: If authentication is successful, you will be logged into the system.
[0890] Step 2:
[0891] Starting a Scenario Session
[0892] Input: User clicks the "Start Scenario Session" button.
[0893] Operation: The terminal calls the voice recognition AI and virtual field module to prepare for the scenario session.
[0894] Data processing and calculation: When preparation is complete, the server generates a ready message.
[0895] Output: The terminal displays the message "Virtual site is ready, start now."
[0896] Step 3:
[0897] Scenario implementation
[0898] Input: The user speaks the on-site response procedures into the virtual scene.
[0899] How it works: The device records the user's voice and sends it to the speech recognition AI and emotion engine in real time.
[0900] Data processing and calculation: Voice data is converted into text data, and then an emotion engine performs emotion analysis.
[0901] Output: Text data and emotional state information are generated.
[0902] Step 4:
[0903] Suggestion content scoring and sentiment analysis
[0904] Input: Text data and emotional state information generated in step 3.
[0905] How it works: The server scores the content of on-site responses based on keywords and phrases extracted from the voice data. The emotion engine analyzes the emotional state.
[0906] Data processing and calculation: Logic, accuracy of information, and appropriate response are scored according to the evaluation criteria, and emotional state is also reflected in the scoring.
[0907] Output: Scoring results and sentiment analysis results are generated.
[0908] Step 5:
[0909] Generate feedback
[0910] Input: The scoring and sentiment analysis results generated in step 4.
[0911] How it works: The server generates feedback based on the score of the suggestions and the results of sentiment analysis.
[0912] Data processing and calculation: Create specific evaluations and advice for improvement for users.
[0913] Output: A feedback message is generated and sent to the terminal.
[0914] Step 6:
[0915] View Feedback
[0916] Input: The feedback message generated in step 5.
[0917] Action: The device displays a feedback message.
[0918] Output: The user sees the evaluation results and advice for improvement.
[0919] Step 7:
[0920] Training mode available
[0921] Input: The feedback message the user confirmed in step 6.
[0922] How it works: Identifies the user's weak areas and offers training modes.
[0923] Data processing and calculation: The server generates strengthening training content based on the user's weaknesses. In some cases, it also adds advice on relaxation.
[0924] Output: Training mode is initiated and specific instructions and advice are displayed to the user.
[0925] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0926] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0927] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0928] [Third embodiment]
[0929] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0930] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0931] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0932] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0933] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0934] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0935] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0936] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0937] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0938] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0939] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0940] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0941] This invention is a system in which a user makes a product proposal by voice, the proposal is analyzed using a voice recognition AI, and an evaluation and feedback are provided. A specific embodiment of this system is described below.
[0942] Basic system configuration
[0943] 1. User logs into the app:
[0944] The user accesses the launched application, enters their ID and password on the login screen, and logs in. The server verifies the entered authentication information and permits the login.
[0945] Example: Salesperson A logs in to the app by entering his ID and password.
[0946] 2. Role-playing session begins:
[0947] When the user clicks the "Start Role Play Session" button, the device calls the voice recognition AI and virtual customer module to prepare for the role play session. After preparation is complete, the device prompts the user to start the session.
[0948] Example: When Sales Representative A clicks the "Start Role Play Session" button, the message "Your virtual customer is ready. Please begin" is displayed.
[0949] 3. Role-playing:
[0950] The user makes product proposals to virtual customers by voice. The device records the user's voice and sends it to a voice recognition AI in real time. The server analyzes the voice data and extracts important keywords and phrases.
[0951] Example: When sales representative A says to a virtual customer, "Hello, today I'd like to introduce you to new product X," the audio is recorded and sent to a speech recognition AI.
[0952] 4. Scoring of proposals:
[0953] The server scores the product proposals based on keywords extracted from the voice data. This score is evaluated based on criteria such as the proposal's logic, the amount of information provided, and the degree to which it meets customer needs. Furthermore, the evaluated score is compared with the scores of other users.
[0954] Example: The server extracts keywords such as "New Product X," "Features," and "Benefits," and based on these, assigns an evaluation score of 50 points. The scores of other sales representatives are 60 points, which is relatively low.
[0955] 5. Feedback and training:
[0956] The server evaluates the suggestions and provides specific feedback to the user, including suggestions for improvement. The device displays this feedback to the user. It also provides specific training modes for users to practice repeatedly. This allows users to focus on areas where they are weak, similar to a TOEIC app.
[0957] Example: The server gives feedback that "your proposal is not logical enough," and the device displays this feedback. In addition, the device is presented with the option to start training to improve logical proposal methods.
[0958] System Applications
[0959] This system allows users to effectively improve their product proposal skills. In particular, new employees can acquire high-level skills in a short period of time by repeatedly receiving standardized feedback and training. This will improve the efficiency and results of sales activities.
[0960] As described above, the present invention allows a user to make a product proposal by voice, and the content of the proposal is analyzed and evaluated, thereby improving product proposal skills and making training of new employees more efficient.
[0961] The processing flow will be explained below.
[0962] Step 1:
[0963] The user launches the app, enters their ID and password on the login screen, and clicks the login button.
[0964] Step 2:
[0965] The terminal sends the authentication information (ID and password) entered by the user to the server.
[0966] Step 3:
[0967] The server checks the transmitted authentication information against the database, and if it is correct, generates session information and sends an authentication success message to the terminal.
[0968] Step 4:
[0969] The device receives the authentication success message and displays the home screen.
[0970] Step 5:
[0971] The user clicks the "Start Role Play Session" button on the home screen.
[0972] Step 6:
[0973] The terminal loads the voice recognition AI and virtual customer module, prepares for the role-playing session, and displays a ready message to the user.
[0974] Step 7:
[0975] The user makes product proposals to the virtual customer by voice.
[0976] Step 8:
[0977] The device records the user's voice in real time and sends the voice data to the voice recognition AI.
[0978] Step 9:
[0979] The server uses voice recognition AI to convert the voice data into text data and extract important keywords and phrases.
[0980] Step 10:
[0981] The server then scores the product proposals based on the analyzed text data. Evaluation criteria include logic, information richness, and suitability to customer needs.
[0982] Step 11:
[0983] The server compares the generated score with the scores of other users and calculates the relativity of the rating.
[0984] Step 12:
[0985] The server generates feedback based on the score and evaluation results and sends the feedback to the device.
[0986] Step 13:
[0987] The terminal displays the feedback received from the server to the user, including logical suggestions for improvement and specific instructions.
[0988] Step 14:
[0989] The device identifies the user's weak points and provides a training mode to improve them.
[0990] Step 15:
[0991] The user selects the training mode and again proposes products to the virtual customer.
[0992] Step 16:
[0993] The device will re-record the audio during training and send it to the voice recognition AI.
[0994] Step 17:
[0995] The server analyzes the new audio data, scores it, and generates feedback.
[0996] Step 18:
[0997] The device will display new feedback to the user, who can then choose to train again or start a new role-playing session.
[0998] This system allows users to continuously improve their product proposal skills.
[0999] Example 1
[1000] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1001] Currently, there are insufficient methods for effectively and efficiently improving the product proposal skills required for sales activities. In particular, while there are systems that evaluate and provide feedback on the logic of proposals, the richness of information, and the degree to which they meet customer needs, there are no concrete methods for streamlining the education and training of new sales representatives. This makes it difficult to improve skills in a short period of time, resulting in problems such as not maximizing the efficiency and results of sales activities.
[1002] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1003] In this invention, the server includes: a means for a user to input voice input to propose a product; a means for recording the voice in real time; a means for transmitting the recorded voice data to a voice recognition AI; a means for analyzing the voice data and extracting important keywords and phrases; a means for scoring the proposal based on the analysis results; a means for comparing the score with that of other users; a means for generating feedback based on the comparison results and presenting it to the user; a means for providing a training mode to strengthen weak areas; a means for invoking the voice recognition AI and a virtual customer module and starting a role-playing session; and a means for notifying the user of the evaluation results. This enables effective and efficient improvement of product proposal skills. In particular, it streamlines the education and training of new sales representatives, allowing them to acquire high-level skills in a short period of time.
[1004] "User" refers to the entity that uses the system to propose products and provide education and training.
[1005] "Merchandise" refers to the products or services that are the subject of sale or proposal.
[1006] "Speech recognition AI" refers to a technology or system that analyzes voice data and extracts important keywords and phrases.
[1007] "Virtual Customer Module" refers to software or a system that simulates a virtual customer with whom a user interacts when making a product proposal in a role-play session.
[1008] "Real time" refers to a state in which processing occurs immediately.
[1009] "Scoring" refers to evaluating the content of product proposals and expressing them in numbers or ranks.
[1010] "Feedback" refers to providing specific improvements and advice to the user's product proposal based on the evaluation results.
[1011] "Training Mode" refers to a practice environment designed to help a user improve a particular skill or knowledge.
[1012] A "role-play session" refers to a practice session in which a user simulates making a product proposal to a virtual customer.
[1013] "Evaluation results" refers to the scores and ranks of the scored proposals, as well as the analysis and comments based on them.
[1014] "Important keywords and phrases" refer to the main words and phrases extracted by speech recognition AI from voice data that are necessary for evaluating product proposals.
[1015] This invention is a system in which a user makes a product proposal by voice, the proposal is analyzed using a voice recognition AI, and an evaluation and feedback are provided. The basic configuration of this system is as follows.
[1016] System Configuration
[1017] 1. A user logs into the app
[1018] The user starts the application on the device and enters their ID and password on the login screen. The device sends the user's input to the server, which then verifies the authentication information. If the information matches, the user is permitted to log in and a login success message is displayed on the device.
[1019] Example: A sales representative enters the ID "user123" and password, presses the login button, and is successfully logged in.
[1020] 2. Start of the role-playing session
[1021] When the user clicks the "Start Role-Play Session" button, the device starts the voice recognition AI and virtual customer module. The server notifies the device that the voice recognition AI and virtual customer module are ready, and the device displays instructions to the user to start the session.
[1022] Example: A salesperson presses the "Start Role Play Session" button, and the terminal displays "Your virtual customer is ready. Begin."
[1023] 3. Role-playing
[1024] The user makes product proposals to the virtual customer by voice. The device records the user's voice and sends it to a voice recognition AI in real time. The server analyzes the voice data and extracts important keywords and phrases.
[1025] Example: When a salesperson says, "Hello, today I'd like to introduce you to our new product," the audio is recorded and sent to a voice recognition AI in real time.
[1026] 4. Scoring of proposals
[1027] The server scores the product proposals based on the extracted keywords and phrases. This score is evaluated based on the proposal's logic, the amount of information provided, and the degree to which it meets customer needs. The score is then compared with that of other users, and the results are recorded.
[1028] Example: The server scores the product based on the keywords "new product," "features," and "benefits," and assigns it a score of 50. This is a relatively low score compared to other users' scores (e.g., 60 points).
[1029] 5. Feedback and training
[1030] Based on the evaluation results, the server generates specific feedback and suggestions for improvement for the user. The device displays the generated feedback to the user and provides a training mode for repeated practice. In this training mode, the user can focus on areas where they are weak.
[1031] Example: The server gives feedback that "your proposal lacks logic," and the device displays this information. In addition, the device displays an option to provide training to improve logical proposal methods.
[1032] Hardware and software used
[1033] Device: The computer, smartphone, tablet, etc. that the user uses.
[1034] Server: Cloud server or on-premise server for voice recognition AI and data analysis.
[1035] Speech Recognition AI: An artificial intelligence model that analyzes a user's speech and extracts important keywords and phrases.
[1036] Virtual Customer Module: A software module that acts as a virtual customer.
[1037] Examples of natural language prompts
[1038] "Hello, I'm a sales representative from X Company. Today I'd like to introduce you to our new product. This product has the following features. It also has the following benefits:
[1039] As described above, this invention allows a user to make a product proposal by voice, and the content of the proposal is analyzed and evaluated, thereby improving product proposal skills and making new employee training more efficient.
[1040] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1041] Step 1:
[1042] Input: The user accesses the app's login screen using their ID and password.
[1043] Operation: The user enters their ID and password and presses the login button.
[1044] Processing: The terminal receives the login information and sends it to the server, which checks the received ID and password against the authentication information in its database.
[1045] Output: The server sends an authentication success message to the terminal, and the terminal displays a login success message to the user.
[1046] Step 2:
[1047] Input: User clicks the "Start Role Play Session" button.
[1048] Operation: The device activates the voice recognition AI and virtual customer module and sends a request to the server.
[1049] Processing: The server receives the request, initializes the speech recognition AI and virtual customer module, and prepares the necessary resources.
[1050] Output: The server sends a ready message to the terminal, and the terminal notifies the user: "The virtual customer is ready. Start now."
[1051] Step 3:
[1052] Input: The user initiates a product proposal by voice.
[1053] Action: A user says to a hypothetical customer, "Hello, today I'd like to introduce you to our new product."
[1054] Processing: The device records the user's voice and sends it to the voice recognition AI in real time.
[1055] Output: The audio data is sent to a server, which analyzes it and extracts important keywords and phrases.
[1056] Step 4:
[1057] Input: Voice data analysis results by voice recognition AI.
[1058] How it works: The server receives the analysis results and scores the suggestions based on keywords and phrases.
[1059] Processing: The server generates a score based on the evaluation criteria (logic of the proposal, richness of information, and degree of adaptation to customer needs).
[1060] Output: The generated score is compared with the scores of other users and the results are stored in a database.
[1061] Step 5:
[1062] Input: Scoring and comparison results.
[1063] Action: The server generates feedback based on the comparison results.
[1064] Processing: The server generates specific improvements and advice as feedback and sends it to the device.
[1065] Output: The device displays feedback to the user and provides a training mode for repeated practice.
[1066] Step 6:
[1067] Input: The user initiates the training mode presented.
[1068] Action: The user selects training mode and begins practicing.
[1069] Processing: The device executes the training mode based on the user's selection and sends the user's progress to the server.
[1070] Output: The server tracks the user's progress and again supports training by providing assessment and feedback.
[1071] (Application example 1)
[1072] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1073] In today's brick-and-mortar stores, the ability of salespeople to effectively propose products to customers is extremely important for increasing sales and customer satisfaction. However, it is not easy, especially for new salespeople, to effectively improve their product proposal skills in a short period of time, and there are few standardized training methods. Under these circumstances, a system is needed to efficiently improve salespeople's product proposal skills.
[1074] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1075] In this invention, the server includes means for a user to input product proposals by voice, means for recording the voice in real time, means for transmitting the recorded voice data to a voice recognition AI, means for analyzing the voice data and extracting important keywords and phrases, means for scoring the proposal content based on the analysis results, means for comparing the score with the scores of other users, means for generating feedback based on the comparison results and presenting it to the user, means for providing a training mode for improving weak areas, and means for displaying the feedback to the salesperson on a smartphone. This allows the salesperson to practice product proposals using voice and receive evaluations and feedback in real time.
[1076] "User" refers to an individual who uses the system to train in product proposals.
[1077] "Product proposal" refers to the act of introducing the features and benefits of a product or service to a customer.
[1078] "Voice input means" refers to a device such as a microphone that allows a user to make product suggestions by voice.
[1079] "Recording means" refers to hardware and software for capturing and storing a user's voice.
[1080] "Voice data" refers to the user's recorded voice information.
[1081] "Voice recognition AI" is an artificial intelligence technology that converts voice data into text data and analyzes the content.
[1082] "Means of analysis" refers to using voice recognition AI to extract important keywords and phrases from voice data.
[1083] "Means for scoring proposal content" refers to evaluating the quality of product proposals and converting them into scores.
[1084] "Means for comparing with scores of other users" refers to a function for comparing the scores of proposals made by multiple users.
[1085] "Feedback" refers to information that indicates the evaluation results and areas for improvement regarding the user's product proposal.
[1086] "Training mode" refers to a mode that provides special functionality that allows users to practice their product proposal skills.
[1087] A "smartphone" is a mobile device that has functions such as voice recording, data transmission and reception, and feedback display.
[1088] This invention relates to a smart assistant application that enables sales staff in brick-and-mortar stores to effectively propose products to customers. This system uses voice recognition AI and feedback functions to train sales staff on how to propose products.
[1089] System Configuration
[1090] The basic configuration of the system is as follows:
[1091] 1. User voice input: The user inputs product proposals by voice. A microphone is used for this purpose.
[1092] 2. Recording of voice data: The device (smartphone) records the user's voice in real time and saves it as voice data.
[1093] 3. Sending to speech recognition AI: The recorded voice data is sent from the device to the speech recognition AI, which converts the voice data into text data and analyzes the content.
[1094] 4. Data analysis: The server uses voice recognition AI to analyze the text data and extract important keywords and phrases.
[1095] 5. Scoring of proposal content: The server evaluates and scores the content of the product proposal based on the extracted keywords and phrases.
[1096] 6. Score comparison: The server compares the scores of multiple users and stores the results.
[1097] 7. Feedback generation and presentation: The server generates feedback based on the score comparison results and presents it to the user. This feedback is displayed on the terminal.
[1098] 8. Providing a training mode: The server provides a training mode to help users improve their weak areas. In this mode, specific instructions and examples are provided.
[1099] Program processing
[1100] Hardware:
[1101] Smartphone: Voice recording, data transmission and reception, feedback display
[1102] software:
[1103] Python: Overall program implementation
[1104] speech_recognition library: Audio recording and speech recognition
[1105] some_ai_module: AI module that evaluates product proposal text
[1106] feedback_module: A module that generates feedback based on the evaluation results.
[1107] Specific examples
[1108] For example, when a salesperson proposes new product Y to a customer, the app records the voice and converts it into text using speech recognition AI. The AI module evaluates the quality of the proposal and provides feedback on the proposal skill based on the evaluation.
[1109] Example prompt for a generative AI model:
[1110] Prompt: "You can be more specific about how product Y compares to other products. You can also explain in detail how it meets your customer's needs."
[1111] This invention allows salespeople to practice product proposals using voice and receive evaluations and feedback in real time, which is expected to improve their proposal skills in a short period of time and contribute to increased customer satisfaction.
[1112] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1113] Step 1:
[1114] Users log in to the app and begin voice input of product proposals. As users speak, they describe the features and benefits of the product, and voice data is generated.
[1115] Input: User's voice
[1116] Output: Audio data
[1117] Specific action: The user speaks into the smartphone's microphone.
[1118] Step 2:
[1119] The device records the user's voice in real time and temporarily stores the recorded voice data.
[1120] Input: Audio data
[1121] Output: Recorded audio data
[1122] Specific operation: The recording function in the smartphone is activated and audio data is captured.
[1123] Step 3:
[1124] The device sends the recorded voice data to the voice recognition AI, which converts the voice data into text data.
[1125] Input: Recorded audio data
[1126] Output: Text data
[1127] How it works: The device sends voice data to a cloud-based voice recognition AI service, which analyzes the voice and converts it into text.
[1128] Step 4:
[1129] The server analyzes the text data sent by the voice recognition AI, extracts important keywords and phrases, and stores the results.
[1130] Input: Text data
[1131] Output: Extracted keywords and phrases
[1132] What it does: Speech recognition AI analyzes the text and picks out important keywords related to your business.
[1133] Step 5:
[1134] The server evaluates and scores the product proposals based on the extracted keywords and phrases, and saves the scoring results.
[1135] Input: Extracted keywords or phrases
[1136] Output: Scoring result (evaluation score)
[1137] What it does: An internal server evaluation algorithm scores the usefulness and persuasiveness of keywords.
[1138] Step 6:
[1139] The server compares the scoring results with the scores of other users and stores the results.
[1140] Input: Scoring results
[1141] Output: Comparison result
[1142] What it does: The server compares the current user's score with the scores of other users in its historical database.
[1143] Step 7:
[1144] The server generates feedback based on the score comparison result and sends the feedback to the terminal, where it is displayed.
[1145] Input: Comparison result
[1146] Output: Feedback
[1147] What it does: Based on the comparison results, the feedback generation module will write down in detail what the user needs to improve and what they should praise. The details will be displayed on the smartphone.
[1148] Step 8:
[1149] A training mode is provided to help users improve their weak points through user input, with specific instructions and examples provided for users to practice again and again.
[1150] Input: Feedback
[1151] Output: Improvement and training plans
[1152] What it does: The training mode starts, providing guidance and examples to help users improve their suggestion skills.
[1153] An example of this prompt would be:
[1154] Example prompt sentence:
[1155] Prompt: "You can be more specific about how product Y compares to other products. You can also explain in detail how it meets your customer's needs."
[1156] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1157] This invention is a system in which users make product proposals by voice, the contents of the proposal are analyzed using voice recognition AI, and the user's emotional state is also analyzed using an emotion engine, providing detailed feedback and a training mode to improve the user's product proposal skills.
[1158] Basic system configuration
[1159] 1. User logs into the app:
[1160] The user accesses the launched application, enters their ID and password on the login screen, and logs in. The server verifies the entered authentication information and permits the login.
[1161] Example: Salesperson A logs in to the app by entering his ID and password.
[1162] 2. Role-playing session begins:
[1163] When the user clicks the "Start Role Play Session" button, the terminal calls the voice recognition AI and virtual customer module, prepares for the role play session, and displays a message to the user that the session is ready.
[1164] Example: When Sales Representative A clicks the "Start Role Play Session" button, the message "Your virtual customer is ready. Please begin" is displayed.
[1165] 3. Role-playing:
[1166] The user makes product proposals to virtual customers by voice, and the device records the user's voice and transmits it in real time to the voice recognition AI and emotion engine.
[1167] Example: When salesperson A says to a virtual customer, "Hello, today I'd like to introduce you to our new product X," the audio is recorded and sent to a speech recognition AI and emotion engine.
[1168] 4. Suggestion scoring and sentiment analysis:
[1169] The server scores the product proposals based on keywords and phrases extracted from the voice data, using criteria such as logic, information richness, and suitability to customer needs.
[1170] The emotion engine analyzes the user's voice data to recognize their emotional state, and the results of this emotion analysis are reflected in the feedback.
[1171] Example: The server extracts keywords such as "New Product X," "Features," and "Benefits," and based on these, assigns a rating of 50 points. The emotion engine detects the user's stress and tension.
[1172] 5. Feedback and training:
[1173] The server generates feedback for the user based on the score of the proposal and the results of the sentiment analysis, and sends it to the device. The device then displays the feedback to the user. The feedback includes not only the logic, information richness, and responsiveness of the proposal to customer needs, but also advice based on the user's emotional state.
[1174] Example: Feedback such as "Your proposal lacks logic" is given, followed by emotional advice such as "We saw some tension, so let's try a different approach."
[1175] 6. Training mode available:
[1176] The device identifies the user's weak areas and offers training modes to improve them. If the user is emotionally unstable, specific advice on how to relax is also added.
[1177] Example: A training mode is initiated where the user can practice logical suggestions, and advice such as "take repeated deep breaths" and "speak confidently" is displayed to further relax.
[1178] System Applications
[1179] This system not only improves users' product proposal skills, but also allows them to control their emotions during presentations. This allows for more effective proposals and improved sales results. New employees, in particular, can acquire advanced skills in a short period of time thanks to consistent feedback and detailed training modes.
[1180] As described above, this invention allows users to make product proposals verbally, analyzes the content of the proposal, and provides feedback and training that takes into account their emotional state, thereby improving product proposal skills and making new employee training more efficient.
[1181] The processing flow will be explained below.
[1182] Step 1:
[1183] The user launches the app, enters their ID and password on the login screen, and clicks the login button.
[1184] Step 2:
[1185] The terminal sends the authentication information (ID and password) entered by the user to the server.
[1186] Step 3:
[1187] The server checks the transmitted authentication information against the database, and if it is correct, generates session information and sends an authentication success message to the terminal.
[1188] Step 4:
[1189] The device receives the authentication success message and displays the home screen.
[1190] Step 5:
[1191] The user clicks the "Start Role Play Session" button on the home screen.
[1192] Step 6:
[1193] The terminal loads the voice recognition AI and virtual customer module, prepares for the role-playing session, and displays a ready message to the user.
[1194] Step 7:
[1195] The user makes product proposals to the virtual customer by voice.
[1196] Step 8:
[1197] The device records the user's voice in real time and sends the voice data to a voice recognition AI and emotion engine.
[1198] Step 9:
[1199] The server uses voice recognition AI to convert the voice data into text data and extract important keywords and phrases.
[1200] Step 10:
[1201] The server uses an emotion engine to analyze the user's emotional state from the voice data.
[1202] Step 11:
[1203] The server then scores the product proposals based on the analyzed text data. Evaluation criteria include logic, information richness, and suitability to customer needs.
[1204] Step 12:
[1205] The server compares the generated score with the scores of other users and calculates the relativity of the rating.
[1206] Step 13:
[1207] The server generates feedback based on the score and the results of the sentiment analysis.
[1208] Step 14:
[1209] The server transmits the generated feedback to the terminal.
[1210] Step 15:
[1211] The terminal displays the feedback received from the server to the user, which includes evaluation results regarding logic, richness of information, and suitability to customer needs, as well as advice based on the user's emotional state.
[1212] Step 16:
[1213] The device offers specific training modes based on the user's weaknesses.
[1214] Step 17:
[1215] The user selects the training mode and again proposes products to the virtual customer.
[1216] Step 18:
[1217] The device records the audio during training and sends it to the voice recognition AI and emotion engine.
[1218] Step 19:
[1219] The server analyzes the new audio data and performs scoring and sentiment analysis.
[1220] Step 20:
[1221] The server again generates feedback and sends it to the device.
[1222] Step 21:
[1223] The device will display new feedback to the user, who can then choose to train again or start a new role-playing session.
[1224] This system allows users to continuously improve their product proposal skills.
[1225] Example 2
[1226] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1227] Conventional product proposal training systems focus on analyzing users' proposal content and providing feedback. However, due to a lack of feedback and training based on the user's emotional state, they do not adequately improve presentation skills or emotional control abilities when proposing products. As a result, many users feel nervous and stressed when proposing products, which reduces the effectiveness of their proposals.
[1228] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a user to input a product proposal by voice; a means for recording the voice in real time; a means for transmitting the recorded voice data to a voice recognition AI; a means for analyzing the voice data and extracting important keywords and phrases; a means for scoring the proposal content based on the analysis results; a means for comparing the score with the scores of other users; a means for generating feedback based on the comparison results and presenting it to the user; a means for transmitting the voice data to an emotion engine and analyzing the emotional state; a means for reflecting the emotion analysis results in the feedback; a means for providing feedback including advice taking the emotional state into consideration; and a means for providing a training mode to strengthen weak areas. This makes it possible to improve not only the content of the product proposal but also the user's ability to control their emotions during the presentation.
[1229] "User" refers to a person or entity who uses the system to train in product proposals.
[1230] "Product proposal" refers to the act of explaining and proposing the features and benefits of a product or service.
[1231] "Voice input means" refers to a method by which a user provides information to a system using speech.
[1232] "Real-time recording means" refers to techniques or methods for instantly recording the voices spoken by a user.
[1233] "Voice recognition AI" refers to artificial intelligence technology that analyzes input voice and converts it into text data.
[1234] An "emotion engine" refers to algorithms and technologies that analyze a user's emotional state from input voice data.
[1235] "Keyword and phrase extraction means" refers to technology that selects important words and expressions from recorded audio data.
[1236] "Means for scoring proposal content" refers to a method for evaluating the quality of product proposals based on the analysis results and assigning a score.
[1237] "Feedback" refers to the comments and advice provided to users based on analyzed data and scores.
[1238] "Training mode" refers to a function or state in which the system provides training to improve the user's weak areas.
[1239] "Advice that takes into account the user's emotional state" refers to advice provided based on the results of an analysis of the user's emotions.
[1240] "Means of comparison" refers to a method for comparing a user's proposal score with the scores of other users.
[1241] This invention is a system that analyzes the content of product proposals made by users through voice and the emotional state of the users. The system aims to improve users' product proposal skills by providing detailed feedback and training modes using voice recognition AI and an emotion engine.
[1242] First, the user logs in to the application. The user enters their ID and password, which are then authenticated by the server. This process uses a device such as a smartphone or PC, and uses an authentication system such as OAuth or LDAP on the server side.
[1243] Next, the user clicks the "Start Role Play Session" button, which causes the system to launch the voice recognition AI (e.g., Google Cloud Speech-to-Text) and virtual customer module. This prepares the role play session, and the device displays to the user, "The virtual customer is ready. Please begin."
[1244] When a user makes a product proposal by voice, the device records this voice and sends it in real time to a voice recognition AI and emotion engine (e.g., IBM Watson Tone Analyzer). Specifically, when a user says, "Hello, today I'd like to introduce you to new product X," the voice is recorded and analyzed.
[1245] The server scores the proposals based on keywords and phrases extracted from the voice data. Evaluation criteria include logic, information richness, and adaptability to customer needs. The server also uses an emotion engine to analyze the user's emotional state, and these results are reflected in the feedback. For example, keywords such as "New Product X," "Features," and "Benefits" are extracted from the voice recognition AI, and a score of 50 is assigned. At the same time, the emotion engine detects the user's state of tension.
[1246] Based on the results of these analyses, the server generates feedback and sends it to the device. The feedback includes the logic of the proposal, the amount of information provided, the degree to which it meets the customer's needs, and emotional advice. Specifically, messages such as "The proposal lacks logic" and "You seem tense, so relax" are displayed on the device.
[1247] The device then identifies the user's weak points and provides a training mode to strengthen them. The user starts the training mode and performs exercises to improve their skills. For example, the device displays a message saying, "You have started a training mode to practice logical proposals," and includes advice such as "Take deep breaths repeatedly" and "Speak with confidence."
[1248] Examples of prompts include "Please enter your ID and password to log in," "Click the button to start the role-playing session," and "Hello, today I'd like to introduce you to our new product X."
[1249] As described above, this invention enables users to improve not only the content of their product proposals but also their emotional control during presentations. This system provides an effective method for improving user skills and streamlining new employee training.
[1250] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1251] Step 1:
[1252] A user logs in to the app.
[1253] Input: The user enters their ID and password.
[1254] Processing: The device sends this authentication information to the server, which verifies it using an authentication system (e.g., OAuth or LDAP). The user's account information is checked against a database.
[1255] Output: If authentication is successful, the user is logged in.
[1256] Specific operation: The user opens the app on their smartphone and enters their ID and password on the login screen. This is then sent from the device to the server, where it is authenticated and, if successful, the user is allowed to log in.
[1257] Step 2:
[1258] The user clicks the "Start Role Play Session" button.
[1259] Input: User clicks the "Start Role Play Session" button.
[1260] Processing: The device calls the speech recognition AI (e.g., Google Cloud Speech-to-Text) and the virtual customer module to prepare for the session.
[1261] Output: The terminal displays "Your virtual customer is ready. Start now."
[1262] Specific operation: When the user clicks the "Start role-playing session" button, the device launches the voice recognition AI and virtual customer module and displays a message on the screen indicating that it is ready.
[1263] Step 3:
[1264] The user makes product suggestions by voice.
[1265] Input: The user makes a product suggestion by voice.
[1266] Processing: The device records the audio and sends it in real time to a speech recognition AI and emotion engine (e.g., IBM Watson Tone Analyzer). The speech recognition AI converts the audio data into text, and the emotion engine analyzes the emotional state.
[1267] Output: Text converted from audio data and sentiment analysis results.
[1268] What it does: When a user says, "Hello, today I'd like to introduce you to new product X," the voice is recorded and the voice data is used for analysis.
[1269] Step 4:
[1270] Analyze the proposal and emotional state.
[1271] Input: Text data from the speech recognition AI and analysis results from the emotion engine.
[1272] Processing: The server scores the suggestions based on keywords and phrases extracted from the speech data, and the emotion engine identifies the emotional state and retrieves the results.
[1273] Output: Suggestion score and sentiment analysis results.
[1274] Specific operation: The server calculates a score based on data obtained from the voice recognition AI (e.g., "New Product X," "Features," and "Benefits"), and the emotion engine identifies emotions such as tension.
[1275] Step 5:
[1276] Provide feedback.
[1277] Input: Suggestion score and sentiment analysis results.
[1278] Processing: The server generates feedback based on this data and sends it to the device. The feedback includes advice based on logic, information richness, responsiveness to customer needs, and emotional state.
[1279] Output: The feedback message.
[1280] Specific behavior: Feedback such as "Your proposal lacks logic" or "You seem nervous, so please relax" will be displayed on the device.
[1281] Step 6:
[1282] Training mode will be implemented.
[1283] Input: Analysis results identifying weak areas and feedback on emotional state.
[1284] Processing: The device offers a training mode to help users improve their weaknesses and also displays specific advice on their emotional state.
[1285] Output: Start of training mode and the accompanying screen display.
[1286] Specific actions: The device will display specific advice such as "You have started a training mode to practice logical suggestions," "Take repeated deep breaths," and "Speak with confidence."
[1287] (Application example 2)
[1288] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1289] Improving on-site response skills is extremely important in modern security services. However, traditional training methods have struggled to effectively support the improvement of individual response abilities and emotional control. In particular, new security guards lack on-site response experience, posing challenges for their performance in situations that require a quick and appropriate response. This calls for improved on-site response quality and more efficient training methods, and a system that solves this problem is needed.
[1290] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1291] In this invention, the server includes: means for a user to input product proposals by voice; means for recording the voice in real time; means for transmitting the recorded voice data to a voice recognition AI; means for analyzing the voice data and extracting important keywords and phrases; means for scoring the proposals based on the analysis results; means for comparing the score with those of other users; means for generating feedback based on the comparison results and presenting it to the user; means for providing a training mode to strengthen weak areas; means for a guard to input on-site response procedures by voice and analyze the contents of the input; and means for providing a training mode to improve on-site response skills based on the analyzed contents. This allows security guards to effectively train their on-site response capabilities and improve their skills, including emotional control.
[1292] "User" refers to the person who operates the system.
[1293] "Merchandise" refers to the products or services that are the subject of sale or proposal.
[1294] "Speech recognition AI" refers to an artificial intelligence algorithm that converts voice data into text data.
[1295] An "emotion engine" refers to software for analyzing a speaker's emotional state from voice data.
[1296] "Virtual customer module" refers to the part of the computer simulation that responds to user suggestions.
[1297] "Scoring" refers to the process of evaluating proposal content and assigning it a score.
[1298] "Feedback" refers to information that provides an evaluation of a user's statements or actions and advice for improvement.
[1299] "Training mode" refers to a practice mode that a user performs to improve a particular skill.
[1300] A "security guard" is a security guard who works to ensure safety on-site.
[1301] "Scene response procedures" refer to a set of steps that outline what a security guard should do in a particular situation.
[1302] "Analysis" refers to the process of analyzing data in detail and extracting useful information.
[1303] "Skills" refer to the techniques and abilities required to perform a particular task or job.
[1304] A specific system for implementing the present invention is a training system whose main purpose is to help security guards improve their on-site response skills. Details of this system are described below.
[1305] Basic system configuration
[1306] 1. Login function:
[1307] The security guard accesses the launched application and enters the ID and password on the login screen. The server verifies the entered authentication information and allows the login.
[1308] Example: A security guard logs into a system by entering an ID and password.
[1309] 2. Begin the scenario session:
[1310] When the security guard clicks the "Start Scenario Session" button, the terminal calls the voice recognition AI and virtual scene module to prepare for the scenario session, and displays a message to the security guard that the session is ready.
[1311] Example: When a security guard clicks the "Start Scenario Session" button, the message "The virtual scene is ready. Please begin" appears.
[1312] 3. Scenario implementation:
[1313] The security guard will then verbally explain the on-site response procedures to the virtual scene, and the device will record the security guard's voice and transmit it to the voice recognition AI and emotion engine in real time.
[1314] Example: When a security guard says to a virtual scene, "I have spotted a suspicious person. I will leave immediately and call the police," the audio is recorded and sent to a speech recognition AI and emotion engine.
[1315] 4. Suggestion scoring and sentiment analysis:
[1316] The server then scores the response based on keywords and phrases extracted from the voice data, with evaluation criteria including logic, accuracy of information, and appropriateness of the response.
[1317] The emotion engine analyzes the security guard's voice data to recognize their emotional state, and the results of this emotion analysis are also reflected in the feedback.
[1318] Example: The server extracts keywords such as "suspicious person," "police," and "report," and based on that, assigns a rating of 70 points. The emotion engine detects the calmness of the security guard.
[1319] 5. Feedback and training:
[1320] The server generates feedback for the security guard based on the score of the suggestion and the results of the emotion analysis, and sends it to the terminal. The terminal then displays the feedback to the security guard. The feedback includes not only the logic of the response, the accuracy of the information, and the appropriateness of the response, but also advice based on the security guard's emotional state.
[1321] Example: Feedback such as "The logic of your response was appropriate" is given, along with emotional advice such as "You showed some composure, so keep it up."
[1322] 6. Training mode available:
[1323] The device identifies weak areas of the guard and offers a training mode to improve them. If the guard is emotionally unstable, specific advice on how to relax is also added.
[1324] Example: A security guard enters a training mode to practice how to call the police, and offers advice such as "take repeated deep breaths" and "calmly assess the situation" to help the guard relax.
[1325] Hardware and software used
[1326] Hardware: Smartphone, head-mounted display
[1327] Software: SpeechRecognition library, Transformer model
[1328] Prompt Sentence Examples
[1329] What you said: A suspicious person has been spotted. Please leave immediately and call the police.
[1330] Emotional state: POSITIVE
[1331] Expected feedback: Your response was logical and appropriate. You showed some composure, so keep it up.
[1332] The above is the details of the specific embodiment for carrying out the present invention.
[1333] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1334] Step 1:
[1335] Log in
[1336] Input: The user enters the ID and password.
[1337] Operation: The device sends the entered ID and password to the server.
[1338] Data processing and calculation: The server checks the authentication information and compares it with the database.
[1339] Output: If authentication is successful, you will be logged into the system.
[1340] Step 2:
[1341] Starting a Scenario Session
[1342] Input: User clicks the "Start Scenario Session" button.
[1343] Operation: The terminal calls the voice recognition AI and virtual field module to prepare for the scenario session.
[1344] Data processing and calculation: When preparation is complete, the server generates a ready message.
[1345] Output: The terminal displays the message "Virtual site is ready, start now."
[1346] Step 3:
[1347] Scenario implementation
[1348] Input: The user speaks the on-site response procedures into the virtual scene.
[1349] How it works: The device records the user's voice and sends it to the speech recognition AI and emotion engine in real time.
[1350] Data processing and calculation: Voice data is converted into text data, and then an emotion engine performs emotion analysis.
[1351] Output: Text data and emotional state information are generated.
[1352] Step 4:
[1353] Suggestion content scoring and sentiment analysis
[1354] Input: Text data and emotional state information generated in step 3.
[1355] How it works: The server scores the content of on-site responses based on keywords and phrases extracted from the voice data. The emotion engine analyzes the emotional state.
[1356] Data processing and calculation: Logic, accuracy of information, and appropriate response are scored according to the evaluation criteria, and emotional state is also reflected in the scoring.
[1357] Output: Scoring results and sentiment analysis results are generated.
[1358] Step 5:
[1359] Generate feedback
[1360] Input: The scoring and sentiment analysis results generated in step 4.
[1361] How it works: The server generates feedback based on the score of the suggestions and the results of sentiment analysis.
[1362] Data processing and calculation: Create specific evaluations and advice for improvement for users.
[1363] Output: A feedback message is generated and sent to the terminal.
[1364] Step 6:
[1365] View Feedback
[1366] Input: The feedback message generated in step 5.
[1367] Action: The device displays a feedback message.
[1368] Output: The user sees the evaluation results and advice for improvement.
[1369] Step 7:
[1370] Training mode available
[1371] Input: The feedback message the user confirmed in step 6.
[1372] How it works: Identifies the user's weak areas and offers training modes.
[1373] Data processing and calculation: The server generates strengthening training content based on the user's weaknesses. In some cases, it also adds advice on relaxation.
[1374] Output: Training mode is initiated and specific instructions and advice are displayed to the user.
[1375] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1376] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1377] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1378] [Fourth embodiment]
[1379] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1380] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1381] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1382] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1383] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1384] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1385] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1386] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1387] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1388] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1389] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1390] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1391] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1392] This invention is a system in which a user makes a product proposal by voice, the proposal is analyzed using a voice recognition AI, and an evaluation and feedback are provided. A specific embodiment of this system is described below.
[1393] Basic system configuration
[1394] 1. User logs into the app:
[1395] The user accesses the launched application, enters their ID and password on the login screen, and logs in. The server verifies the entered authentication information and permits the login.
[1396] Example: Salesperson A logs in to the app by entering his ID and password.
[1397] 2. Role-playing session begins:
[1398] When the user clicks the "Start Role Play Session" button, the device calls the voice recognition AI and virtual customer module to prepare for the role play session. After preparation is complete, the device prompts the user to start the session.
[1399] Example: When Sales Representative A clicks the "Start Role Play Session" button, the message "Your virtual customer is ready. Please begin" is displayed.
[1400] 3. Role-playing:
[1401] The user makes product proposals to virtual customers by voice. The device records the user's voice and sends it to a voice recognition AI in real time. The server analyzes the voice data and extracts important keywords and phrases.
[1402] Example: When sales representative A says to a virtual customer, "Hello, today I'd like to introduce you to new product X," the audio is recorded and sent to a speech recognition AI.
[1403] 4. Scoring of proposals:
[1404] The server scores the product proposals based on keywords extracted from the voice data. This score is evaluated based on criteria such as the proposal's logic, the amount of information provided, and the degree to which it meets customer needs. Furthermore, the evaluated score is compared with the scores of other users.
[1405] Example: The server extracts keywords such as "New Product X," "Features," and "Benefits," and based on these, assigns an evaluation score of 50 points. The scores of other sales representatives are 60 points, which is relatively low.
[1406] 5. Feedback and training:
[1407] The server evaluates the suggestions and provides specific feedback to the user, including suggestions for improvement. The device displays this feedback to the user. It also provides specific training modes for users to practice repeatedly. This allows users to focus on areas where they are weak, similar to a TOEIC app.
[1408] Example: The server gives feedback that "your proposal is not logical enough," and the device displays this feedback. In addition, the device is presented with the option to start training to improve logical proposal methods.
[1409] System Applications
[1410] This system allows users to effectively improve their product proposal skills. In particular, new employees can acquire high-level skills in a short period of time by repeatedly receiving standardized feedback and training. This will improve the efficiency and results of sales activities.
[1411] As described above, the present invention allows a user to make a product proposal by voice, and the content of the proposal is analyzed and evaluated, thereby improving product proposal skills and making training of new employees more efficient.
[1412] The processing flow will be explained below.
[1413] Step 1:
[1414] The user launches the app, enters their ID and password on the login screen, and clicks the login button.
[1415] Step 2:
[1416] The terminal sends the authentication information (ID and password) entered by the user to the server.
[1417] Step 3:
[1418] The server checks the transmitted authentication information against the database, and if it is correct, generates session information and sends an authentication success message to the terminal.
[1419] Step 4:
[1420] The device receives the authentication success message and displays the home screen.
[1421] Step 5:
[1422] The user clicks the "Start Role Play Session" button on the home screen.
[1423] Step 6:
[1424] The terminal loads the voice recognition AI and virtual customer module, prepares for the role-playing session, and displays a ready message to the user.
[1425] Step 7:
[1426] The user makes product proposals to the virtual customer by voice.
[1427] Step 8:
[1428] The device records the user's voice in real time and sends the voice data to the voice recognition AI.
[1429] Step 9:
[1430] The server uses voice recognition AI to convert the voice data into text data and extract important keywords and phrases.
[1431] Step 10:
[1432] The server then scores the product proposals based on the analyzed text data. Evaluation criteria include logic, information richness, and suitability to customer needs.
[1433] Step 11:
[1434] The server compares the generated score with the scores of other users and calculates the relativity of the rating.
[1435] Step 12:
[1436] The server generates feedback based on the score and evaluation results and sends the feedback to the device.
[1437] Step 13:
[1438] The terminal displays the feedback received from the server to the user, including logical suggestions for improvement and specific instructions.
[1439] Step 14:
[1440] The device identifies the user's weak points and provides a training mode to improve them.
[1441] Step 15:
[1442] The user selects the training mode and again proposes products to the virtual customer.
[1443] Step 16:
[1444] The device will re-record the audio during training and send it to the voice recognition AI.
[1445] Step 17:
[1446] The server analyzes the new audio data, scores it, and generates feedback.
[1447] Step 18:
[1448] The device will display new feedback to the user, who can then choose to train again or start a new role-playing session.
[1449] This system allows users to continuously improve their product proposal skills.
[1450] Example 1
[1451] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1452] Currently, there are insufficient methods for effectively and efficiently improving the product proposal skills required for sales activities. In particular, while there are systems that evaluate and provide feedback on the logic of proposals, the richness of information, and the degree to which they meet customer needs, there are no concrete methods for streamlining the education and training of new sales representatives. This makes it difficult to improve skills in a short period of time, resulting in problems such as not maximizing the efficiency and results of sales activities.
[1453] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1454] In this invention, the server includes: a means for a user to input voice input to propose a product; a means for recording the voice in real time; a means for transmitting the recorded voice data to a voice recognition AI; a means for analyzing the voice data and extracting important keywords and phrases; a means for scoring the proposal based on the analysis results; a means for comparing the score with that of other users; a means for generating feedback based on the comparison results and presenting it to the user; a means for providing a training mode to strengthen weak areas; a means for invoking the voice recognition AI and a virtual customer module and starting a role-playing session; and a means for notifying the user of the evaluation results. This enables effective and efficient improvement of product proposal skills. In particular, it streamlines the education and training of new sales representatives, allowing them to acquire high-level skills in a short period of time.
[1455] "User" refers to the entity that uses the system to propose products and provide education and training.
[1456] "Merchandise" refers to the products or services that are the subject of sale or proposal.
[1457] "Speech recognition AI" refers to a technology or system that analyzes voice data and extracts important keywords and phrases.
[1458] "Virtual Customer Module" refers to software or a system that simulates a virtual customer with whom a user interacts when making a product proposal in a role-play session.
[1459] "Real time" refers to a state in which processing occurs immediately.
[1460] "Scoring" refers to evaluating the content of product proposals and expressing them in numbers or ranks.
[1461] "Feedback" refers to providing specific improvements and advice to the user's product proposal based on the evaluation results.
[1462] "Training Mode" refers to a practice environment designed to help a user improve a particular skill or knowledge.
[1463] A "role-play session" refers to a practice session in which a user simulates making a product proposal to a virtual customer.
[1464] "Evaluation results" refers to the scores and ranks of the scored proposals, as well as the analysis and comments based on them.
[1465] "Important keywords and phrases" refer to the main words and phrases extracted by speech recognition AI from voice data that are necessary for evaluating product proposals.
[1466] This invention is a system in which a user makes a product proposal by voice, the proposal is analyzed using a voice recognition AI, and an evaluation and feedback are provided. The basic configuration of this system is as follows.
[1467] System Configuration
[1468] 1. A user logs into the app
[1469] The user starts the application on the device and enters their ID and password on the login screen. The device sends the user's input to the server, which then verifies the authentication information. If the information matches, the user is permitted to log in and a login success message is displayed on the device.
[1470] Example: A sales representative enters the ID "user123" and password, presses the login button, and is successfully logged in.
[1471] 2. Start of the role-playing session
[1472] When the user clicks the "Start Role-Play Session" button, the device starts the voice recognition AI and virtual customer module. The server notifies the device that the voice recognition AI and virtual customer module are ready, and the device displays instructions to the user to start the session.
[1473] Example: A salesperson presses the "Start Role Play Session" button, and the terminal displays "Your virtual customer is ready. Begin."
[1474] 3. Role-playing
[1475] The user makes product proposals to the virtual customer by voice. The device records the user's voice and sends it to a voice recognition AI in real time. The server analyzes the voice data and extracts important keywords and phrases.
[1476] Example: When a salesperson says, "Hello, today I'd like to introduce you to our new product," the audio is recorded and sent to a voice recognition AI in real time.
[1477] 4. Scoring of proposals
[1478] The server scores the product proposals based on the extracted keywords and phrases. This score is evaluated based on the proposal's logic, the amount of information provided, and the degree to which it meets customer needs. The score is then compared with that of other users, and the results are recorded.
[1479] Example: The server scores the product based on the keywords "new product," "features," and "benefits," and assigns it a score of 50. This is a relatively low score compared to other users' scores (e.g., 60 points).
[1480] 5. Feedback and training
[1481] Based on the evaluation results, the server generates specific feedback and suggestions for improvement for the user. The device displays the generated feedback to the user and provides a training mode for repeated practice. In this training mode, the user can focus on areas where they are weak.
[1482] Example: The server gives feedback that "your proposal lacks logic," and the device displays this information. In addition, the device displays an option to provide training to improve logical proposal methods.
[1483] Hardware and software used
[1484] Device: The computer, smartphone, tablet, etc. that the user uses.
[1485] Server: Cloud server or on-premise server for voice recognition AI and data analysis.
[1486] Speech Recognition AI: An artificial intelligence model that analyzes a user's speech and extracts important keywords and phrases.
[1487] Virtual Customer Module: A software module that acts as a virtual customer.
[1488] Examples of natural language prompts
[1489] "Hello, I'm a sales representative from X Company. Today I'd like to introduce you to our new product. This product has the following features. It also has the following benefits:
[1490] As described above, this invention allows a user to make a product proposal by voice, and the content of the proposal is analyzed and evaluated, thereby improving product proposal skills and making new employee training more efficient.
[1491] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1492] Step 1:
[1493] Input: The user accesses the app's login screen using their ID and password.
[1494] Operation: The user enters their ID and password and presses the login button.
[1495] Processing: The terminal receives the login information and sends it to the server, which checks the received ID and password against the authentication information in its database.
[1496] Output: The server sends an authentication success message to the terminal, and the terminal displays a login success message to the user.
[1497] Step 2:
[1498] Input: User clicks the "Start Role Play Session" button.
[1499] Operation: The device activates the voice recognition AI and virtual customer module and sends a request to the server.
[1500] Processing: The server receives the request, initializes the speech recognition AI and virtual customer module, and prepares the necessary resources.
[1501] Output: The server sends a ready message to the terminal, and the terminal notifies the user: "The virtual customer is ready. Start now."
[1502] Step 3:
[1503] Input: The user initiates a product proposal by voice.
[1504] Action: A user says to a hypothetical customer, "Hello, today I'd like to introduce you to our new product."
[1505] Processing: The device records the user's voice and sends it to the voice recognition AI in real time.
[1506] Output: The audio data is sent to a server, which analyzes it and extracts important keywords and phrases.
[1507] Step 4:
[1508] Input: Voice data analysis results by voice recognition AI.
[1509] How it works: The server receives the analysis results and scores the suggestions based on keywords and phrases.
[1510] Processing: The server generates a score based on the evaluation criteria (logic of the proposal, richness of information, and degree of adaptation to customer needs).
[1511] Output: The generated score is compared with the scores of other users and the results are stored in a database.
[1512] Step 5:
[1513] Input: Scoring and comparison results.
[1514] Action: The server generates feedback based on the comparison results.
[1515] Processing: The server generates specific improvements and advice as feedback and sends it to the device.
[1516] Output: The device displays feedback to the user and provides a training mode for repeated practice.
[1517] Step 6:
[1518] Input: The user initiates the training mode presented.
[1519] Action: The user selects training mode and begins practicing.
[1520] Processing: The device executes the training mode based on the user's selection and sends the user's progress to the server.
[1521] Output: The server tracks the user's progress and again supports training by providing assessment and feedback.
[1522] (Application example 1)
[1523] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1524] In today's brick-and-mortar stores, the ability of salespeople to effectively propose products to customers is extremely important for increasing sales and customer satisfaction. However, it is not easy, especially for new salespeople, to effectively improve their product proposal skills in a short period of time, and there are few standardized training methods. Under these circumstances, a system is needed to efficiently improve salespeople's product proposal skills.
[1525] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1526] In this invention, the server includes means for a user to input product proposals by voice, means for recording the voice in real time, means for transmitting the recorded voice data to a voice recognition AI, means for analyzing the voice data and extracting important keywords and phrases, means for scoring the proposal content based on the analysis results, means for comparing the score with the scores of other users, means for generating feedback based on the comparison results and presenting it to the user, means for providing a training mode for improving weak areas, and means for displaying the feedback to the salesperson on a smartphone. This allows the salesperson to practice product proposals using voice and receive evaluations and feedback in real time.
[1527] "User" refers to an individual who uses the system to train in product proposals.
[1528] "Product proposal" refers to the act of introducing the features and benefits of a product or service to a customer.
[1529] "Voice input means" refers to a device such as a microphone that allows a user to make product suggestions by voice.
[1530] "Recording means" refers to hardware and software for capturing and storing a user's voice.
[1531] "Voice data" refers to the user's recorded voice information.
[1532] "Voice recognition AI" is an artificial intelligence technology that converts voice data into text data and analyzes the content.
[1533] "Means of analysis" refers to using voice recognition AI to extract important keywords and phrases from voice data.
[1534] "Means for scoring proposal content" refers to evaluating the quality of product proposals and converting them into scores.
[1535] "Means for comparing with scores of other users" refers to a function for comparing the scores of proposals made by multiple users.
[1536] "Feedback" refers to information that indicates the evaluation results and areas for improvement regarding the user's product proposal.
[1537] "Training mode" refers to a mode that provides special functionality that allows users to practice their product proposal skills.
[1538] A "smartphone" is a mobile device that has functions such as voice recording, data transmission and reception, and feedback display.
[1539] This invention relates to a smart assistant application that enables sales staff in brick-and-mortar stores to effectively propose products to customers. This system uses voice recognition AI and feedback functions to train sales staff on how to propose products.
[1540] System Configuration
[1541] The basic configuration of the system is as follows:
[1542] 1. User voice input: The user inputs product proposals by voice. A microphone is used for this purpose.
[1543] 2. Recording of voice data: The device (smartphone) records the user's voice in real time and saves it as voice data.
[1544] 3. Sending to speech recognition AI: The recorded voice data is sent from the device to the speech recognition AI, which converts the voice data into text data and analyzes the content.
[1545] 4. Data analysis: The server uses voice recognition AI to analyze the text data and extract important keywords and phrases.
[1546] 5. Scoring of proposal content: The server evaluates and scores the content of the product proposal based on the extracted keywords and phrases.
[1547] 6. Score comparison: The server compares the scores of multiple users and stores the results.
[1548] 7. Feedback generation and presentation: The server generates feedback based on the score comparison results and presents it to the user. This feedback is displayed on the terminal.
[1549] 8. Providing a training mode: The server provides a training mode to help users improve their weak areas. In this mode, specific instructions and examples are provided.
[1550] Program processing
[1551] Hardware:
[1552] Smartphone: Voice recording, data transmission and reception, feedback display
[1553] software:
[1554] Python: Overall program implementation
[1555] speech_recognition library: Audio recording and speech recognition
[1556] some_ai_module: AI module that evaluates product proposal text
[1557] feedback_module: A module that generates feedback based on the evaluation results.
[1558] Specific examples
[1559] For example, when a salesperson proposes new product Y to a customer, the app records the voice and converts it into text using speech recognition AI. The AI module evaluates the quality of the proposal and provides feedback on the proposal skill based on the evaluation.
[1560] Example prompt for a generative AI model:
[1561] Prompt: "You can be more specific about how product Y compares to other products. You can also explain in detail how it meets your customer's needs."
[1562] This invention allows salespeople to practice product proposals using voice and receive evaluations and feedback in real time, which is expected to improve their proposal skills in a short period of time and contribute to increased customer satisfaction.
[1563] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1564] Step 1:
[1565] Users log in to the app and begin voice input of product proposals. As users speak, they describe the features and benefits of the product, and voice data is generated.
[1566] Input: User's voice
[1567] Output: Audio data
[1568] Specific action: The user speaks into the smartphone's microphone.
[1569] Step 2:
[1570] The device records the user's voice in real time and temporarily stores the recorded voice data.
[1571] Input: Audio data
[1572] Output: Recorded audio data
[1573] Specific operation: The recording function in the smartphone is activated and audio data is captured.
[1574] Step 3:
[1575] The device sends the recorded voice data to the voice recognition AI, which converts the voice data into text data.
[1576] Input: Recorded audio data
[1577] Output: Text data
[1578] How it works: The device sends voice data to a cloud-based voice recognition AI service, which analyzes the voice and converts it into text.
[1579] Step 4:
[1580] The server analyzes the text data sent by the voice recognition AI, extracts important keywords and phrases, and stores the results.
[1581] Input: Text data
[1582] Output: Extracted keywords and phrases
[1583] What it does: Speech recognition AI analyzes the text and picks out important keywords related to your business.
[1584] Step 5:
[1585] The server evaluates and scores the product proposals based on the extracted keywords and phrases, and saves the scoring results.
[1586] Input: Extracted keywords or phrases
[1587] Output: Scoring result (evaluation score)
[1588] What it does: An internal server evaluation algorithm scores the usefulness and persuasiveness of keywords.
[1589] Step 6:
[1590] The server compares the scoring results with the scores of other users and stores the results.
[1591] Input: Scoring results
[1592] Output: Comparison result
[1593] What it does: The server compares the current user's score with the scores of other users in its historical database.
[1594] Step 7:
[1595] The server generates feedback based on the score comparison result and sends the feedback to the terminal, where it is displayed.
[1596] Input: Comparison result
[1597] Output: Feedback
[1598] What it does: Based on the comparison results, the feedback generation module will write down in detail what the user needs to improve and what they should praise. The details will be displayed on the smartphone.
[1599] Step 8:
[1600] A training mode is provided to help users improve their weak points through user input, with specific instructions and examples provided for users to practice again and again.
[1601] Input: Feedback
[1602] Output: Improvement and training plans
[1603] What it does: The training mode starts, providing guidance and examples to help users improve their suggestion skills.
[1604] An example of this prompt would be:
[1605] Example prompt sentence:
[1606] Prompt: "You can be more specific about how product Y compares to other products. You can also explain in detail how it meets your customer's needs."
[1607] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1608] This invention is a system in which users make product proposals by voice, the contents of the proposal are analyzed using voice recognition AI, and the user's emotional state is also analyzed using an emotion engine, providing detailed feedback and a training mode to improve the user's product proposal skills.
[1609] Basic system configuration
[1610] 1. User logs into the app:
[1611] The user accesses the launched application, enters their ID and password on the login screen, and logs in. The server verifies the entered authentication information and permits the login.
[1612] Example: Salesperson A logs in to the app by entering his ID and password.
[1613] 2. Role-playing session begins:
[1614] When the user clicks the "Start Role Play Session" button, the terminal calls the voice recognition AI and virtual customer module, prepares for the role play session, and displays a message to the user that the session is ready.
[1615] Example: When Sales Representative A clicks the "Start Role Play Session" button, the message "Your virtual customer is ready. Please begin" is displayed.
[1616] 3. Role-playing:
[1617] The user makes product proposals to virtual customers by voice, and the device records the user's voice and transmits it in real time to the voice recognition AI and emotion engine.
[1618] Example: When salesperson A says to a virtual customer, "Hello, today I'd like to introduce you to our new product X," the audio is recorded and sent to a speech recognition AI and emotion engine.
[1619] 4. Suggestion scoring and sentiment analysis:
[1620] The server scores the product proposals based on keywords and phrases extracted from the voice data, using criteria such as logic, information richness, and suitability to customer needs.
[1621] The emotion engine analyzes the user's voice data to recognize their emotional state, and the results of this emotion analysis are reflected in the feedback.
[1622] Example: The server extracts keywords such as "New Product X," "Features," and "Benefits," and based on these, assigns a rating of 50 points. The emotion engine detects the user's stress and tension.
[1623] 5. Feedback and training:
[1624] The server generates feedback for the user based on the score of the proposal and the results of the sentiment analysis, and sends it to the device. The device then displays the feedback to the user. The feedback includes not only the logic, information richness, and responsiveness of the proposal to customer needs, but also advice based on the user's emotional state.
[1625] Example: Feedback such as "Your proposal lacks logic" is given, followed by emotional advice such as "We saw some tension, so let's try a different approach."
[1626] 6. Training mode available:
[1627] The device identifies the user's weak areas and offers training modes to improve them. If the user is emotionally unstable, specific advice on how to relax is also added.
[1628] Example: A training mode is initiated where the user can practice logical suggestions, and advice such as "take repeated deep breaths" and "speak confidently" is displayed to further relax.
[1629] System Applications
[1630] This system not only improves users' product proposal skills, but also allows them to control their emotions during presentations. This allows for more effective proposals and improved sales results. New employees, in particular, can acquire advanced skills in a short period of time thanks to consistent feedback and detailed training modes.
[1631] As described above, this invention allows users to make product proposals verbally, analyzes the content of the proposal, and provides feedback and training that takes into account their emotional state, thereby improving product proposal skills and making new employee training more efficient.
[1632] The processing flow will be explained below.
[1633] Step 1:
[1634] The user launches the app, enters their ID and password on the login screen, and clicks the login button.
[1635] Step 2:
[1636] The terminal sends the authentication information (ID and password) entered by the user to the server.
[1637] Step 3:
[1638] The server checks the transmitted authentication information against the database, and if it is correct, generates session information and sends an authentication success message to the terminal.
[1639] Step 4:
[1640] The device receives the authentication success message and displays the home screen.
[1641] Step 5:
[1642] The user clicks the "Start Role Play Session" button on the home screen.
[1643] Step 6:
[1644] The terminal loads the voice recognition AI and virtual customer module, prepares for the role-playing session, and displays a ready message to the user.
[1645] Step 7:
[1646] The user makes product proposals to the virtual customer by voice.
[1647] Step 8:
[1648] The device records the user's voice in real time and sends the voice data to a voice recognition AI and emotion engine.
[1649] Step 9:
[1650] The server uses voice recognition AI to convert the voice data into text data and extract important keywords and phrases.
[1651] Step 10:
[1652] The server uses an emotion engine to analyze the user's emotional state from the voice data.
[1653] Step 11:
[1654] The server then scores the product proposals based on the analyzed text data. Evaluation criteria include logic, information richness, and suitability to customer needs.
[1655] Step 12:
[1656] The server compares the generated score with the scores of other users and calculates the relativity of the rating.
[1657] Step 13:
[1658] The server generates feedback based on the score and the results of the sentiment analysis.
[1659] Step 14:
[1660] The server transmits the generated feedback to the terminal.
[1661] Step 15:
[1662] The terminal displays the feedback received from the server to the user, which includes evaluation results regarding logic, richness of information, and suitability to customer needs, as well as advice based on the user's emotional state.
[1663] Step 16:
[1664] The device offers specific training modes based on the user's weaknesses.
[1665] Step 17:
[1666] The user selects the training mode and again proposes products to the virtual customer.
[1667] Step 18:
[1668] The device records the audio during training and sends it to the voice recognition AI and emotion engine.
[1669] Step 19:
[1670] The server analyzes the new audio data and performs scoring and sentiment analysis.
[1671] Step 20:
[1672] The server again generates feedback and sends it to the device.
[1673] Step 21:
[1674] The device will display new feedback to the user, who can then choose to train again or start a new role-playing session.
[1675] This system allows users to continuously improve their product proposal skills.
[1676] Example 2
[1677] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1678] Conventional product proposal training systems focus on analyzing users' proposal content and providing feedback. However, due to a lack of feedback and training based on the user's emotional state, they do not adequately improve presentation skills or emotional control abilities when proposing products. As a result, many users feel nervous and stressed when proposing products, which reduces the effectiveness of their proposals.
[1679] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a user to input a product proposal by voice; a means for recording the voice in real time; a means for transmitting the recorded voice data to a voice recognition AI; a means for analyzing the voice data and extracting important keywords and phrases; a means for scoring the proposal content based on the analysis results; a means for comparing the score with the scores of other users; a means for generating feedback based on the comparison results and presenting it to the user; a means for transmitting the voice data to an emotion engine and analyzing the emotional state; a means for reflecting the emotion analysis results in the feedback; a means for providing feedback including advice taking the emotional state into consideration; and a means for providing a training mode to strengthen weak areas. This makes it possible to improve not only the content of the product proposal but also the user's ability to control their emotions during the presentation.
[1680] "User" refers to a person or entity who uses the system to train in product proposals.
[1681] "Product proposal" refers to the act of explaining and proposing the features and benefits of a product or service.
[1682] "Voice input means" refers to a method by which a user provides information to a system using speech.
[1683] "Real-time recording means" refers to techniques or methods for instantly recording the voices spoken by a user.
[1684] "Voice recognition AI" refers to artificial intelligence technology that analyzes input voice and converts it into text data.
[1685] An "emotion engine" refers to algorithms and technologies that analyze a user's emotional state from input voice data.
[1686] "Keyword and phrase extraction means" refers to technology that selects important words and expressions from recorded audio data.
[1687] "Means for scoring proposal content" refers to a method for evaluating the quality of product proposals based on the analysis results and assigning a score.
[1688] "Feedback" refers to the comments and advice provided to users based on analyzed data and scores.
[1689] "Training mode" refers to a function or state in which the system provides training to improve the user's weak areas.
[1690] "Advice that takes into account the user's emotional state" refers to advice provided based on the results of an analysis of the user's emotions.
[1691] "Means of comparison" refers to a method for comparing a user's proposal score with the scores of other users.
[1692] This invention is a system that analyzes the content of product proposals made by users through voice and the emotional state of the users. The system aims to improve users' product proposal skills by providing detailed feedback and training modes using voice recognition AI and an emotion engine.
[1693] First, the user logs in to the application. The user enters their ID and password, which are then authenticated by the server. This process uses a device such as a smartphone or PC, and uses an authentication system such as OAuth or LDAP on the server side.
[1694] Next, the user clicks the "Start Role Play Session" button, which causes the system to launch the voice recognition AI (e.g., Google Cloud Speech-to-Text) and virtual customer module. This prepares the role play session, and the device displays to the user, "The virtual customer is ready. Please begin."
[1695] When a user makes a product proposal by voice, the device records this voice and sends it in real time to a voice recognition AI and emotion engine (e.g., IBM Watson Tone Analyzer). Specifically, when a user says, "Hello, today I'd like to introduce you to new product X," the voice is recorded and analyzed.
[1696] The server scores the proposals based on keywords and phrases extracted from the voice data. Evaluation criteria include logic, information richness, and adaptability to customer needs. The server also uses an emotion engine to analyze the user's emotional state, and these results are reflected in the feedback. For example, keywords such as "New Product X," "Features," and "Benefits" are extracted from the voice recognition AI, and a score of 50 is assigned. At the same time, the emotion engine detects the user's state of tension.
[1697] Based on the results of these analyses, the server generates feedback and sends it to the device. The feedback includes the logic of the proposal, the amount of information provided, the degree to which it meets the customer's needs, and emotional advice. Specifically, messages such as "The proposal lacks logic" and "You seem tense, so relax" are displayed on the device.
[1698] The device then identifies the user's weak points and provides a training mode to strengthen them. The user starts the training mode and performs exercises to improve their skills. For example, the device displays a message saying, "You have started a training mode to practice logical proposals," and includes advice such as "Take deep breaths repeatedly" and "Speak with confidence."
[1699] Examples of prompts include "Please enter your ID and password to log in," "Click the button to start the role-playing session," and "Hello, today I'd like to introduce you to our new product X."
[1700] As described above, this invention enables users to improve not only the content of their product proposals but also their emotional control during presentations. This system provides an effective method for improving user skills and streamlining new employee training.
[1701] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1702] Step 1:
[1703] A user logs in to the app.
[1704] Input: The user enters their ID and password.
[1705] Processing: The device sends this authentication information to the server, which verifies it using an authentication system (e.g., OAuth or LDAP). The user's account information is checked against a database.
[1706] Output: If authentication is successful, the user is logged in.
[1707] Specific operation: The user opens the app on their smartphone and enters their ID and password on the login screen. This is then sent from the device to the server, where it is authenticated and, if successful, the user is allowed to log in.
[1708] Step 2:
[1709] The user clicks the "Start Role Play Session" button.
[1710] Input: User clicks the "Start Role Play Session" button.
[1711] Processing: The device calls the speech recognition AI (e.g., Google Cloud Speech-to-Text) and the virtual customer module to prepare for the session.
[1712] Output: The terminal displays "Your virtual customer is ready. Start now."
[1713] Specific operation: When the user clicks the "Start role-playing session" button, the device launches the voice recognition AI and virtual customer module and displays a message on the screen indicating that it is ready.
[1714] Step 3:
[1715] The user makes product suggestions by voice.
[1716] Input: The user makes a product suggestion by voice.
[1717] Processing: The device records the audio and sends it in real time to a speech recognition AI and emotion engine (e.g., IBM Watson Tone Analyzer). The speech recognition AI converts the audio data into text, and the emotion engine analyzes the emotional state.
[1718] Output: Text converted from audio data and sentiment analysis results.
[1719] What it does: When a user says, "Hello, today I'd like to introduce you to new product X," the voice is recorded and the voice data is used for analysis.
[1720] Step 4:
[1721] Analyze the proposal and emotional state.
[1722] Input: Text data from the speech recognition AI and analysis results from the emotion engine.
[1723] Processing: The server scores the suggestions based on keywords and phrases extracted from the speech data, and the emotion engine identifies the emotional state and retrieves the results.
[1724] Output: Suggestion score and sentiment analysis results.
[1725] Specific operation: The server calculates a score based on data obtained from the voice recognition AI (e.g., "New Product X," "Features," and "Benefits"), and the emotion engine identifies emotions such as tension.
[1726] Step 5:
[1727] Provide feedback.
[1728] Input: Suggestion score and sentiment analysis results.
[1729] Processing: The server generates feedback based on this data and sends it to the device. The feedback includes advice based on logic, information richness, responsiveness to customer needs, and emotional state.
[1730] Output: The feedback message.
[1731] Specific behavior: Feedback such as "Your proposal lacks logic" or "You seem nervous, so please relax" will be displayed on the device.
[1732] Step 6:
[1733] Training mode will be implemented.
[1734] Input: Analysis results identifying weak areas and feedback on emotional state.
[1735] Processing: The device offers a training mode to help users improve their weaknesses and also displays specific advice on their emotional state.
[1736] Output: Start of training mode and the accompanying screen display.
[1737] Specific actions: The device will display specific advice such as "You have started a training mode to practice logical suggestions," "Take repeated deep breaths," and "Speak with confidence."
[1738] (Application example 2)
[1739] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1740] Improving on-site response skills is extremely important in modern security services. However, traditional training methods have struggled to effectively support the improvement of individual response abilities and emotional control. In particular, new security guards lack on-site response experience, posing challenges for their performance in situations that require a quick and appropriate response. This calls for improved on-site response quality and more efficient training methods, and a system that solves this problem is needed.
[1741] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1742] In this invention, the server includes: means for a user to input product proposals by voice; means for recording the voice in real time; means for transmitting the recorded voice data to a voice recognition AI; means for analyzing the voice data and extracting important keywords and phrases; means for scoring the proposals based on the analysis results; means for comparing the score with those of other users; means for generating feedback based on the comparison results and presenting it to the user; means for providing a training mode to strengthen weak areas; means for a guard to input on-site response procedures by voice and analyze the contents of the input; and means for providing a training mode to improve on-site response skills based on the analyzed contents. This allows security guards to effectively train their on-site response capabilities and improve their skills, including emotional control.
[1743] "User" refers to the person who operates the system.
[1744] "Merchandise" refers to the products or services that are the subject of sale or proposal.
[1745] "Speech recognition AI" refers to an artificial intelligence algorithm that converts voice data into text data.
[1746] An "emotion engine" refers to software for analyzing a speaker's emotional state from voice data.
[1747] "Virtual customer module" refers to the part of the computer simulation that responds to user suggestions.
[1748] "Scoring" refers to the process of evaluating proposal content and assigning it a score.
[1749] "Feedback" refers to information that provides an evaluation of a user's statements or actions and advice for improvement.
[1750] "Training mode" refers to a practice mode that a user performs to improve a particular skill.
[1751] A "security guard" is a security guard who works to ensure safety on-site.
[1752] "Scene response procedures" refer to a set of steps that outline what a security guard should do in a particular situation.
[1753] "Analysis" refers to the process of analyzing data in detail and extracting useful information.
[1754] "Skills" refer to the techniques and abilities required to perform a particular task or job.
[1755] A specific system for implementing the present invention is a training system whose main purpose is to help security guards improve their on-site response skills. Details of this system are described below.
[1756] Basic system configuration
[1757] 1. Login function:
[1758] The security guard accesses the launched application and enters the ID and password on the login screen. The server verifies the entered authentication information and allows the login.
[1759] Example: A security guard logs into a system by entering an ID and password.
[1760] 2. Begin the scenario session:
[1761] When the security guard clicks the "Start Scenario Session" button, the terminal calls the voice recognition AI and virtual scene module to prepare for the scenario session, and displays a message to the security guard that the session is ready.
[1762] Example: When a security guard clicks the "Start Scenario Session" button, the message "The virtual scene is ready. Please begin" appears.
[1763] 3. Scenario implementation:
[1764] The security guard will then verbally explain the on-site response procedures to the virtual scene, and the device will record the security guard's voice and transmit it to the voice recognition AI and emotion engine in real time.
[1765] Example: When a security guard says to a virtual scene, "I have spotted a suspicious person. I will leave immediately and call the police," the audio is recorded and sent to a speech recognition AI and emotion engine.
[1766] 4. Suggestion scoring and sentiment analysis:
[1767] The server then scores the response based on keywords and phrases extracted from the voice data, with evaluation criteria including logic, accuracy of information, and appropriateness of the response.
[1768] The emotion engine analyzes the security guard's voice data to recognize their emotional state, and the results of this emotion analysis are also reflected in the feedback.
[1769] Example: The server extracts keywords such as "suspicious person," "police," and "report," and based on that, assigns a rating of 70 points. The emotion engine detects the calmness of the security guard.
[1770] 5. Feedback and training:
[1771] The server generates feedback for the security guard based on the score of the suggestion and the results of the emotion analysis, and sends it to the terminal. The terminal then displays the feedback to the security guard. The feedback includes not only the logic of the response, the accuracy of the information, and the appropriateness of the response, but also advice based on the security guard's emotional state.
[1772] Example: Feedback such as "The logic of your response was appropriate" is given, along with emotional advice such as "You showed some composure, so keep it up."
[1773] 6. Training mode available:
[1774] The device identifies weak areas of the guard and offers a training mode to improve them. If the guard is emotionally unstable, specific advice on how to relax is also added.
[1775] Example: A security guard enters a training mode to practice how to call the police, and offers advice such as "take repeated deep breaths" and "calmly assess the situation" to help the guard relax.
[1776] Hardware and software used
[1777] Hardware: Smartphone, head-mounted display
[1778] Software: SpeechRecognition library, Transformer model
[1779] Prompt Sentence Examples
[1780] What you said: A suspicious person has been spotted. Please leave immediately and call the police.
[1781] Emotional state: POSITIVE
[1782] Expected feedback: Your response was logical and appropriate. You showed some composure, so keep it up.
[1783] The above is the details of the specific embodiment for carrying out the present invention.
[1784] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1785] Step 1:
[1786] Log in
[1787] Input: The user enters the ID and password.
[1788] Operation: The device sends the entered ID and password to the server.
[1789] Data processing and calculation: The server checks the authentication information and compares it with the database.
[1790] Output: If authentication is successful, you will be logged into the system.
[1791] Step 2:
[1792] Starting a Scenario Session
[1793] Input: User clicks the "Start Scenario Session" button.
[1794] Operation: The terminal calls the voice recognition AI and virtual field module to prepare for the scenario session.
[1795] Data processing and calculation: When preparation is complete, the server generates a ready message.
[1796] Output: The terminal displays the message "Virtual site is ready, start now."
[1797] Step 3:
[1798] Scenario implementation
[1799] Input: The user speaks the on-site response procedures into the virtual scene.
[1800] How it works: The device records the user's voice and sends it to the speech recognition AI and emotion engine in real time.
[1801] Data processing and calculation: Voice data is converted into text data, and then an emotion engine performs emotion analysis.
[1802] Output: Text data and emotional state information are generated.
[1803] Step 4:
[1804] Suggestion content scoring and sentiment analysis
[1805] Input: Text data and emotional state information generated in step 3.
[1806] How it works: The server scores the content of on-site responses based on keywords and phrases extracted from the voice data. The emotion engine analyzes the emotional state.
[1807] Data processing and calculation: Logic, accuracy of information, and appropriate response are scored according to the evaluation criteria, and emotional state is also reflected in the scoring.
[1808] Output: Scoring results and sentiment analysis results are generated.
[1809] Step 5:
[1810] Generate feedback
[1811] Input: The scoring and sentiment analysis results generated in step 4.
[1812] How it works: The server generates feedback based on the score of the suggestions and the results of sentiment analysis.
[1813] Data processing and calculation: Create specific evaluations and advice for improvement for users.
[1814] Output: A feedback message is generated and sent to the terminal.
[1815] Step 6:
[1816] View Feedback
[1817] Input: The feedback message generated in step 5.
[1818] Action: The device displays a feedback message.
[1819] Output: The user sees the evaluation results and advice for improvement.
[1820] Step 7:
[1821] Training mode available
[1822] Input: The feedback message the user confirmed in step 6.
[1823] How it works: Identifies the user's weak areas and offers training modes.
[1824] Data processing and calculation: The server generates strengthening training content based on the user's weaknesses. In some cases, it also adds advice on relaxation.
[1825] Output: Training mode is initiated and specific instructions and advice are displayed to the user.
[1826] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1827] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1828] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1829] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1830] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1831] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1832] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1833] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1834] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1835] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1836] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1837] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1838] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1839] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1840] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1841] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1842] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1843] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1844] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1845] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1846] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1847] The following is further disclosed regarding the above embodiment.
[1848] (Claim 1)
[1849] A means for a user to input product suggestions by voice;
[1850] means for recording said audio in real time;
[1851] A means for transmitting the recorded voice data to a voice recognition AI;
[1852] means for analyzing the speech data and extracting important keywords and phrases;
[1853] A means for scoring the proposal content based on the analysis result;
[1854] means for comparing said score with scores of other users;
[1855] means for generating and presenting feedback to the user based on the comparison results;
[1856] A means to provide training modes to strengthen weak areas;
[1857] A system including:
[1858] (Claim 2)
[1859] 2. The system according to claim 1, wherein the content of the feedback includes evaluation results regarding the logic of the proposal, the richness of information, and the degree of response to customer needs.
[1860] (Claim 3)
[1861] 10. The system of claim 1, wherein the training mode includes specific instructions and examples for practicing areas where the user struggles.
[1862] "Example 1"
[1863] (Claim 1)
[1864] a means for a user to input a voice message to suggest a commodity;
[1865] means for recording said audio in real time;
[1866] A means for transmitting the recorded voice data to a voice recognition AI;
[1867] means for analyzing the speech data and extracting important keywords and phrases;
[1868] A means for scoring the proposal content based on the analysis result;
[1869] means for comparing said score with scores of other users;
[1870] means for generating and presenting feedback to the user based on the comparison results;
[1871] A means to provide training modes to strengthen weak areas;
[1872] A means to invoke the voice recognition AI and virtual customer module and initiate role-play sessions;
[1873] a means for notifying the evaluation results;
[1874] A system including:
[1875] (Claim 2)
[1876] 2. The system according to claim 1, wherein the content of the feedback includes evaluation results regarding the logic of the proposal, the richness of information, and the degree of response to customer needs.
[1877] (Claim 3)
[1878] 10. The system of claim 1, wherein the training mode includes specific instructions and examples for practicing areas where the user struggles.
[1879] "Application Example 1"
[1880] (Claim 1)
[1881] A means for a user to input product suggestions by voice;
[1882] means for recording said audio in real time;
[1883] A means for transmitting the recorded voice data to a voice recognition AI;
[1884] means for analyzing the speech data and extracting important keywords and phrases;
[1885] A means for scoring the proposal content based on the analysis result;
[1886] means for comparing said score with scores of other users;
[1887] means for generating and presenting feedback to the user based on the comparison results;
[1888] A means to provide training modes to strengthen weak areas;
[1889] means for displaying said feedback to a salesperson on a smartphone;
[1890] A system including:
[1891] (Claim 2)
[1892] 2. The system according to claim 1, wherein the content of the feedback includes evaluation results regarding the logic of the proposal, the richness of information, and the degree of response to customer needs.
[1893] (Claim 3)
[1894] 10. The system of claim 1, wherein the training mode includes specific instructions and examples for practicing areas where the user struggles.
[1895] "Example 2: Combining Emotion Engines"
[1896] (Claim 1)
[1897] A means for a user to input product suggestions by voice;
[1898] means for recording said audio in real time;
[1899] A means for transmitting the recorded voice data to a voice recognition AI;
[1900] means for analyzing the speech data and extracting important keywords and phrases;
[1901] A means for scoring the proposal content based on the analysis result;
[1902] means for comparing said score with scores of other users;
[1903] means for generating and presenting feedback to the user based on the comparison results;
[1904] means for transmitting the voice data to an emotion engine and analyzing the emotional state;
[1905] a means for reflecting the emotion analysis result in feedback;
[1906] a means for providing feedback including advice taking into account emotional state;
[1907] A means to provide training modes to strengthen weak areas;
[1908] A system including:
[1909] (Claim 2)
[1910] 2. The system according to claim 1, wherein the feedback includes evaluation results regarding the logic of the proposal, the richness of information, and the degree to which the proposal satisfies customer needs, as well as advice based on the user's emotional state.
[1911] (Claim 3)
[1912] 10. The system of claim 1, wherein the training mode includes specific instructions and examples for practicing areas where the user struggles, as well as specific advice regarding emotional states.
[1913] "Application example 2 when combining emotion engines"
[1914] (Claim 1)
[1915] A means for a user to input product suggestions by voice;
[1916] means for recording said audio in real time;
[1917] A means for transmitting the recorded voice data to a voice recognition AI;
[1918] means for analyzing the speech data and extracting important keywords and phrases;
[1919] A means for scoring the proposal content based on the analysis result;
[1920] means for comparing said score with scores of other users;
[1921] means for generating and presenting feedback to the user based on the comparison results;
[1922] A means to provide training modes to strengthen weak areas;
[1923] A means for guards to input on-site response procedures by voice and analyze the contents of those instructions;
[1924] a means for providing a training mode for improving on-site response skills based on the analyzed content;
[1925] A system including:
[1926] (Claim 2)
[1927] 2. The system according to claim 1, wherein the content of the feedback includes evaluation results regarding the logic of the proposal, the richness of information, and the degree of response to customer needs.
[1928] (Claim 3)
[1929] 10. The system of claim 1, wherein the training mode includes specific instructions and examples for practicing areas of weakness for the user, and advice on maintaining composure. [Explanation of symbols]
[1930] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for a user to input product suggestions by voice; means for recording said audio in real time; A means for transmitting the recorded voice data to a voice recognition AI; means for analyzing the speech data and extracting important keywords and phrases; A means for scoring the proposal content based on the analysis result; means for comparing said score with scores of other users; means for generating and presenting feedback to the user based on the comparison results; A means to provide training modes to strengthen weak areas; A system including:
2. The system according to claim 1 , wherein the content of the feedback includes evaluation results regarding the logic of the proposal, the richness of information, and the degree of response to customer needs.
3. The system of claim 1 , wherein the training mode includes specific instructions and examples for practicing areas of difficulty for the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A