system

The AI-based customer harassment training system uses scenario selection, voice and motion generation, and report compilation to create immersive training scenarios, enhancing employee response skills and providing detailed feedback.

JP2026072766APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing systems fail to effectively train employees on customer harassment, lacking comprehensive and realistic simulation tools.

Method used

A fully automated AI system comprising a scenario selection unit, voice generation unit, action generation unit, and report generation unit, utilizing speech synthesis, video generation, and natural language generation technologies to simulate customer interactions and compile training reports.

Benefits of technology

Enables effective and realistic training for customer harassment scenarios, providing detailed feedback and improving employee response capabilities through immersive simulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072766000001_ABST
    Figure 2026072766000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to effectively conduct training on customer harassment. [Solution] The system according to the embodiment comprises a scenario selection unit, a voice generation unit, an action generation unit, a conversation generation unit, and a report generation unit. The scenario selection unit selects a scenario. The voice generation unit generates a customer's voice based on the scenario selected by the scenario selection unit. The action generation unit generates a customer's action based on the scenario selected by the scenario selection unit. The conversation generation unit generates a customer's conversation based on the scenario selected by the scenario selection unit. The report generation unit compiles the results generated by the voice generation unit, action generation unit, and conversation generation unit into a report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, training for customer harassment has not been effectively carried out, and there is room for improvement.

[0005] The system according to the embodiment aims to effectively conduct training for customer harassment.

Means for Solving the Problems

[0006] The system according to this embodiment comprises a scenario selection unit, a voice generation unit, an action generation unit, a conversation generation unit, and a report generation unit. The scenario selection unit selects a scenario. The voice generation unit generates a customer's voice based on the scenario selected by the scenario selection unit. The action generation unit generates a customer's action based on the scenario selected by the scenario selection unit. The conversation generation unit generates a customer's conversation based on the scenario selected by the scenario selection unit. The report generation unit compiles the results generated by the voice generation unit, action generation unit, and conversation generation unit into a report. [Effects of the Invention]

[0007] The system according to this embodiment can effectively conduct training on customer harassment. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) This system is a fully automated AI generation system for conducting training on customer harassment. Specifically, it consists of the following steps: First, the user selects a scenario. Next, the generation AI generates the customer's voice, full body movements, and conversation based on the selected scenario. The voice is generated by speech synthesis AI, the full body movements are realized by video generation AI, and the conversation is generated by the generation AI. When the simulation starts, the user interacts with the generated customer and engages in a conversation. The results of the simulation are compiled into a report that can be reviewed by the supervisor. This system is expected to enable satisfactory training on customer harassment and contribute to the elimination of customer harassment. This allows for the fully automated implementation of training on customer harassment.

[0029] The customer harassment training system according to this embodiment comprises a scenario selection unit, a voice generation unit, an action generation unit, a conversation generation unit, and a report generation unit. The scenario selection unit allows the user to select a scenario. The scenario selection unit can select from multiple scenarios, such as training scenarios and simulation scenarios. The voice generation unit generates a customer's voice based on the selected scenario. The voice generation unit generates a customer's voice using, for example, speech synthesis technology. For example, the voice generation unit can adjust the tone and pitch of the customer's voice using speech synthesis AI. The action generation unit generates a customer's actions based on the selected scenario. The action generation unit generates a full-body action of the customer using, for example, action capture technology. For example, the action generation unit can generate the type and pattern of the customer's actions using video generation AI. The conversation generation unit generates a customer's conversation based on the selected scenario. The conversation generation unit generates a customer's conversation using, for example, natural language generation technology. For example, the conversation generation unit can generate the content and tone of the customer's conversation using generation AI. The report generation unit compiles the results generated by the voice generation unit, action generation unit, and conversation generation unit into a report. The report generation unit can, for example, compile the generated results into a text report or a graphical report. This allows the customer harassment training system according to this embodiment to conduct customer harassment training fully automatically.

[0030] The scenario selection section allows users to choose a scenario. This section offers multiple scenarios, such as training scenarios and simulation scenarios. Specifically, it provides an intuitive interface, displaying detailed descriptions, objectives, and difficulty levels for each scenario. Users can use this information to select the scenario best suited to their training objectives. The scenario selection section also includes a function to recommend the optimal scenario, taking into account past training history and the user's skill level. For example, it automatically suggests the next scenario to challenge based on scenarios the user has previously completed and their results. Furthermore, based on the user's selected scenario, the scenario selection section collaborates with other departments and systems to prepare necessary data and resources. This allows users to smoothly begin training. Additionally, the scenario selection section can collect user feedback and continuously improve the content and difficulty level of the scenarios. This enables the scenario selection section to provide users with an optimal training experience and improve their understanding of and ability to respond to customer harassment.

[0031] The voice generation unit generates customer voices based on selected scenarios. For example, it uses speech synthesis technology to generate customer voices. Specifically, the voice generation unit can adjust the tone and pitch of customer voices using speech synthesis AI. The speech synthesis AI learns from various pre-collected audio data to generate realistic customer voices. For example, if a customer is angry, the voice tone can be raised and the pitch increased to create a sense of tension. Conversely, if a customer is calm, the voice tone can be lowered and the pitch stabilized to create a calm atmosphere. Depending on the scenario, the voice generation unit can also generate the voices of multiple customers simultaneously, simulating situations where multiple customers speak at the same time. Furthermore, the voice generation unit has a function to play the generated audio in real time, allowing users to hear actual customer voices during training. This enables the voice generation unit to provide users with a realistic training environment and improve their ability to respond to customer harassment.

[0032] The motion generation unit generates customer movements based on a selected scenario. For example, it can generate full-body movements of a customer using motion capture technology. Specifically, the motion generation unit can generate types and patterns of customer movements using video generation AI. The video generation AI learns from diverse motion data collected in advance and can generate realistic customer movements. For example, if a customer is angry, it can generate movements such as raising their hands or leaning forward. If a customer is calm, it can generate movements such as crossing their arms or moving slowly. Depending on the scenario, the motion generation unit can also generate movements for multiple customers simultaneously, simulating situations where multiple customers are moving at the same time. Furthermore, the motion generation unit has a function to display the generated movements in real time, allowing users to see customer movements during actual training. This enables the motion generation unit to provide users with a realistic training environment and improve their ability to respond to customer harassment.

[0033] The conversation generation unit generates customer conversations based on selected scenarios. For example, it uses natural language generation technology to generate customer conversations. Specifically, the conversation generation unit can use a generation AI to generate the content and tone of customer conversations. The generation AI learns from diverse conversation data collected in advance and can generate realistic customer conversations. For example, if a customer is angry, it can generate aggressive language and a strong tone. Conversely, if a customer is calm, it can generate polite language and a gentle tone. Depending on the scenario, the conversation generation unit can also generate conversations for multiple customers simultaneously, simulating situations where multiple customers speak at the same time. Furthermore, the conversation generation unit has a function to play back the generated conversations in real time, allowing users to listen to customer conversations during actual training. This enables the conversation generation unit to provide users with a realistic training environment and improve their ability to respond to customer harassment.

[0034] The report generation unit compiles the results generated by the voice generation unit, motion generation unit, and conversation generation unit into a report. For example, the report generation unit can compile the generated results into text reports or graphical reports. Specifically, the report generation unit integrates data provided by each department and analyzes the user's training results in detail. For example, it records what scenarios the user selected, what actions they took, and what feedback they received as a result. Based on this information, the report generation unit generates a report that clearly shows the user's strengths and areas for improvement. Furthermore, the report generation unit can visually represent the user's performance using graphs and charts. This allows the user to grasp their training results at a glance and set specific goals for the next training session. The report generation unit also has the function to output the generated reports in PDF and Excel formats, allowing users to easily save and share them. This enables the report generation unit to effectively analyze the user's training results and provide specific feedback to improve their ability to respond to customer harassment.

[0035] The scenario selection unit can analyze past training history and automatically suggest the most suitable scenario to the user. For example, the scenario selection unit can suggest the next scenario to tackle based on the results of scenarios the user has previously worked on. For example, the scenario selection unit can also suggest scenarios that the user has previously struggled with to encourage them to overcome those difficulties. For example, the scenario selection unit can suggest even more difficult scenarios based on scenarios that the user has previously excelled at. This allows the system to provide the most suitable scenario based on the user's past training history. Some or all of the above-described processes in the scenario selection unit may be performed using AI, for example, or without AI. For example, the scenario selection unit can input the user's past training data into a generating AI and have the generating AI suggest the most suitable scenario.

[0036] The scenario selection unit can customize scenarios based on the user's job duties and experience when a scenario is selected. For example, if the user is a new employee, the scenario selection unit will provide a basic scenario. If the user is a manager, the scenario selection unit can also provide a scenario that allows the user to demonstrate leadership. If the user is skilled in a particular task, the scenario selection unit can also provide a scenario related to that task. This allows the system to provide scenarios tailored to the user's job duties and experience. Some or all of the above-described processes in the scenario selection unit may be performed using AI, for example, or without AI. For example, the scenario selection unit can input the user's job data into a generating AI and have the generating AI perform the scenario customization.

[0037] The voice generation unit can apply different voice qualities depending on the scenario content during voice generation. For example, in a severe scenario, the voice generation unit will generate a severe voice quality. In a gentle scenario, for example, the voice generation unit can also generate a soft voice quality. In a neutral scenario, for example, the voice generation unit can also generate a standard voice quality. This allows for the provision of appropriate voice qualities according to the scenario content. Some or all of the above processing in the voice generation unit may be performed using AI, for example, or without AI. For example, the voice generation unit can input scenario content data into a generation AI and have the generation AI perform the application of voice qualities.

[0038] The motion generation unit can apply different motion patterns depending on the scenario content when generating motions. For example, in the case of a severe scenario, the motion generation unit will generate a strict motion. For example, in the case of a gentle scenario, the motion generation unit can also generate a gentle motion. For example, in the case of a neutral scenario, the motion generation unit can also generate a standard motion. This allows for the provision of appropriate motion patterns according to the scenario content. Some or all of the above-described processing in the motion generation unit may be performed using AI, for example, or without AI. For example, the motion generation unit can input scenario content data into a generation AI and have the generation AI execute the application of motion patterns.

[0039] The conversation generation unit can apply different conversation patterns depending on the scenario content when generating conversations. For example, in the case of a severe scenario, the conversation generation unit will generate a strict conversation pattern. For example, in the case of a gentle scenario, the conversation generation unit can also generate a gentle conversation pattern. For example, in the case of a neutral scenario, the conversation generation unit can also generate a standard conversation pattern. This allows for the provision of appropriate conversation patterns according to the content of the scenario. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can input scenario content data into a generation AI and have the generation AI perform the application of conversation patterns.

[0040] The report generation unit can apply different report formats based on the scenario content and results when generating reports. For example, in the case of a severe scenario, the report generation unit will generate a strict report format. For example, in the case of an easy scenario, the report generation unit can also generate a lenient report format. For example, in the case of a neutral scenario, the report generation unit can also generate a standard report format. This allows for the provision of an appropriate report format according to the scenario content and results. Some or all of the above processing in the report generation unit may be performed using AI, for example, or without AI. For example, the report generation unit can input scenario content data into a generation AI and have the generation AI perform the application of the report format.

[0041] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0042] The customer harassment training system can also include a real-time assistant. The real-time assistant provides real-time advice to the user during simulations. For example, if the user faces a difficult situation, it can suggest appropriate responses. Furthermore, the real-time assistant can analyze the user's past training history and provide advice based on past successes. For instance, it can suggest responses that worked for the user in the past and instruct the user to avoid responses that failed in the past. This allows for the provision of appropriate advice based on the user's past training history.

[0043] The customer harassment training system can also include a scenario generation unit. This unit automatically generates new scenarios based on the user's job duties and experience. For example, if a user is engaged in a specific task, it can generate scenarios related to that task. Furthermore, the scenario generation unit can analyze the user's past training history and generate variations of previously completed scenarios. For instance, it can adjust the difficulty level of a previously completed scenario and offer it as a new scenario. This allows for the provision of new scenarios tailored to the user's job duties and experience.

[0044] The customer harassment training system can also include a data analysis department. This department analyzes user training data and evaluates the effectiveness of the training. For example, it can compare users' performance before and after training and quantify the training's effectiveness. Furthermore, the data analysis department can predict training effectiveness based on users' past training history. For instance, it can predict the effectiveness of the next training a user should receive based on the effectiveness of their past training. This allows for appropriate evaluation and prediction based on users' training data.

[0045] The customer harassment training system can also include a customization settings section. This section provides users with the ability to freely customize the training settings. For example, users can set the difficulty level of scenarios and the duration of the training. Furthermore, the customization settings section can suggest optimal settings based on the user's past training history. For instance, it can adjust the difficulty level of scenarios the user previously struggled with and increase the difficulty level of scenarios the user excelled at. This allows for the provision of appropriate training settings tailored to the user's needs.

[0046] The customer harassment training system can also be equipped with a multilingual support function. This function allows users to receive training in different languages. For example, users can select scenarios in multiple languages, such as English and Chinese. Furthermore, the multilingual support function can automatically translate scenario content and feedback based on the user's language settings. For example, if a user selects Japanese, all scenarios and feedback will be provided in Japanese. This ensures that appropriate training can be provided to users who speak different languages.

[0047] The following briefly describes the processing flow for example form 1.

[0048] Step 1: The scenario selection section allows the user to select a scenario. The scenario selection section allows the user to choose from multiple scenarios, such as a training scenario or a simulation scenario. Step 2: The voice generation unit generates the customer's voice based on the selected scenario. The voice generation unit generates the customer's voice using, for example, speech synthesis technology. For example, the voice generation unit can adjust the tone and pitch of the customer's voice using speech synthesis AI. Step 3: The motion generation unit generates customer movements based on the selected scenario. The motion generation unit generates full-body movements of the customer, for example, using motion capture technology. For example, the motion generation unit can generate types and patterns of customer movements using video generation AI. Step 4: The conversation generation unit generates customer conversations based on the selected scenario. The conversation generation unit generates customer conversations using, for example, natural language generation technology. For example, the conversation generation unit can use generation AI to generate the content and tone of customer conversations. Step 5: The report generation unit compiles the results generated by the voice generation unit, action generation unit, and conversation generation unit into a report. The report generation unit can, for example, compile the generated results into a text report or a graphical report.

[0049] (Example of form 2) This system is a fully automated AI generation system for conducting training on customer harassment. Specifically, it consists of the following steps: First, the user selects a scenario. Next, the generation AI generates the customer's voice, full body movements, and conversation based on the selected scenario. The voice is generated by speech synthesis AI, the full body movements are realized by video generation AI, and the conversation is generated by the generation AI. When the simulation starts, the user interacts with the generated customer and engages in a conversation. The results of the simulation are compiled into a report that can be reviewed by the supervisor. This system is expected to enable satisfactory training on customer harassment and contribute to the elimination of customer harassment. This allows for the fully automated implementation of training on customer harassment.

[0050] The customer harassment training system according to this embodiment comprises a scenario selection unit, a voice generation unit, an action generation unit, a conversation generation unit, and a report generation unit. The scenario selection unit allows the user to select a scenario. The scenario selection unit can select from multiple scenarios, such as training scenarios and simulation scenarios. The voice generation unit generates a customer's voice based on the selected scenario. The voice generation unit generates a customer's voice using, for example, speech synthesis technology. For example, the voice generation unit can adjust the tone and pitch of the customer's voice using speech synthesis AI. The action generation unit generates a customer's actions based on the selected scenario. The action generation unit generates a full-body action of the customer using, for example, action capture technology. For example, the action generation unit can generate the type and pattern of the customer's actions using video generation AI. The conversation generation unit generates a customer's conversation based on the selected scenario. The conversation generation unit generates a customer's conversation using, for example, natural language generation technology. For example, the conversation generation unit can generate the content and tone of the customer's conversation using generation AI. The report generation unit compiles the results generated by the voice generation unit, action generation unit, and conversation generation unit into a report. The report generation unit can, for example, compile the generated results into a text report or a graphical report. This allows the customer harassment training system according to this embodiment to conduct customer harassment training fully automatically.

[0051] The scenario selection section allows users to choose a scenario. This section offers multiple scenarios, such as training scenarios and simulation scenarios. Specifically, it provides an intuitive interface, displaying detailed descriptions, objectives, and difficulty levels for each scenario. Users can use this information to select the scenario best suited to their training objectives. The scenario selection section also includes a function to recommend the optimal scenario, taking into account past training history and the user's skill level. For example, it automatically suggests the next scenario to challenge based on scenarios the user has previously completed and their results. Furthermore, based on the user's selected scenario, the scenario selection section collaborates with other departments and systems to prepare necessary data and resources. This allows users to smoothly begin training. Additionally, the scenario selection section can collect user feedback and continuously improve the content and difficulty level of the scenarios. This enables the scenario selection section to provide users with an optimal training experience and improve their understanding of and ability to respond to customer harassment.

[0052] The voice generation unit generates customer voices based on selected scenarios. For example, it uses speech synthesis technology to generate customer voices. Specifically, the voice generation unit can adjust the tone and pitch of customer voices using speech synthesis AI. The speech synthesis AI learns from various pre-collected audio data to generate realistic customer voices. For example, if a customer is angry, the voice tone can be raised and the pitch increased to create a sense of tension. Conversely, if a customer is calm, the voice tone can be lowered and the pitch stabilized to create a calm atmosphere. Depending on the scenario, the voice generation unit can also generate the voices of multiple customers simultaneously, simulating situations where multiple customers speak at the same time. Furthermore, the voice generation unit has a function to play the generated audio in real time, allowing users to hear actual customer voices during training. This enables the voice generation unit to provide users with a realistic training environment and improve their ability to respond to customer harassment.

[0053] The motion generation unit generates customer movements based on a selected scenario. For example, it can generate full-body movements of a customer using motion capture technology. Specifically, the motion generation unit can generate types and patterns of customer movements using video generation AI. The video generation AI learns from diverse motion data collected in advance and can generate realistic customer movements. For example, if a customer is angry, it can generate movements such as raising their hands or leaning forward. If a customer is calm, it can generate movements such as crossing their arms or moving slowly. Depending on the scenario, the motion generation unit can also generate movements for multiple customers simultaneously, simulating situations where multiple customers are moving at the same time. Furthermore, the motion generation unit has a function to display the generated movements in real time, allowing users to see customer movements during actual training. This enables the motion generation unit to provide users with a realistic training environment and improve their ability to respond to customer harassment.

[0054] The conversation generation unit generates customer conversations based on selected scenarios. For example, it uses natural language generation technology to generate customer conversations. Specifically, the conversation generation unit can use a generation AI to generate the content and tone of customer conversations. The generation AI learns from diverse conversation data collected in advance and can generate realistic customer conversations. For example, if a customer is angry, it can generate aggressive language and a strong tone. Conversely, if a customer is calm, it can generate polite language and a gentle tone. Depending on the scenario, the conversation generation unit can also generate conversations for multiple customers simultaneously, simulating situations where multiple customers speak at the same time. Furthermore, the conversation generation unit has a function to play back the generated conversations in real time, allowing users to listen to customer conversations during actual training. This enables the conversation generation unit to provide users with a realistic training environment and improve their ability to respond to customer harassment.

[0055] The report generation unit compiles the results generated by the voice generation unit, motion generation unit, and conversation generation unit into a report. For example, the report generation unit can compile the generated results into text reports or graphical reports. Specifically, the report generation unit integrates data provided by each department and analyzes the user's training results in detail. For example, it records what scenarios the user selected, what actions they took, and what feedback they received as a result. Based on this information, the report generation unit generates a report that clearly shows the user's strengths and areas for improvement. Furthermore, the report generation unit can visually represent the user's performance using graphs and charts. This allows the user to grasp their training results at a glance and set specific goals for the next training session. The report generation unit also has the function to output the generated reports in PDF and Excel formats, allowing users to easily save and share them. This enables the report generation unit to effectively analyze the user's training results and provide specific feedback to improve their ability to respond to customer harassment.

[0056] The scenario selection unit can estimate the user's emotions and adjust the difficulty of the scenario based on the estimated emotions. For example, if the user is feeling stressed, the scenario selection unit can set the difficulty of the scenario low and provide an easy scenario. For example, if the user is relaxed, the scenario selection unit can also set the difficulty of the scenario high and provide a challenging scenario. For example, if the user has neutral emotions, the scenario selection unit can also set the difficulty of the scenario to medium and provide a balanced scenario. This allows for the provision of scenarios of appropriate difficulty according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the scenario selection unit may be performed using AI, for example, or without AI. For example, the scenario selection unit can input the user's facial expression data into the generative AI and have the generative AI perform emotion estimation.

[0057] The scenario selection unit can analyze past training history and automatically suggest the most suitable scenario to the user. For example, the scenario selection unit can suggest the next scenario to tackle based on the results of scenarios the user has previously worked on. For example, the scenario selection unit can also suggest scenarios that the user has previously struggled with to encourage them to overcome those difficulties. For example, the scenario selection unit can suggest even more difficult scenarios based on scenarios that the user has previously excelled at. This allows the system to provide the most suitable scenario based on the user's past training history. Some or all of the above-described processes in the scenario selection unit may be performed using AI, for example, or without AI. For example, the scenario selection unit can input the user's past training data into a generating AI and have the generating AI suggest the most suitable scenario.

[0058] The scenario selection unit can customize scenarios based on the user's job duties and experience when a scenario is selected. For example, if the user is a new employee, the scenario selection unit will provide a basic scenario. If the user is a manager, the scenario selection unit can also provide a scenario that allows the user to demonstrate leadership. If the user is skilled in a particular task, the scenario selection unit can also provide a scenario related to that task. This allows the system to provide scenarios tailored to the user's job duties and experience. Some or all of the above-described processes in the scenario selection unit may be performed using AI, for example, or without AI. For example, the scenario selection unit can input the user's job data into a generating AI and have the generating AI perform the scenario customization.

[0059] The voice generation unit can estimate the user's emotions and adjust the tone and speed of the voice based on the estimated emotions. For example, if the user is nervous, the voice generation unit can generate a voice in a calm tone and at a slow speed. For example, if the user is relaxed, the voice generation unit can also generate a voice in a bright tone and at a natural speed. For example, if the user is in a hurry, the voice generation unit can also generate a voice in a quick and concise tone. This allows for the provision of an appropriate tone and speed of voice according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the voice generation unit may be performed using AI, for example, or without AI. For example, the voice generation unit can input the user's voice data into the generation AI and have the generation AI adjust the tone and speed of the voice.

[0060] The voice generation unit can apply different voice qualities depending on the scenario content during voice generation. For example, in a severe scenario, the voice generation unit will generate a severe voice quality. In a gentle scenario, for example, the voice generation unit can also generate a soft voice quality. In a neutral scenario, for example, the voice generation unit can also generate a standard voice quality. This allows for the provision of appropriate voice qualities according to the scenario content. Some or all of the above processing in the voice generation unit may be performed using AI, for example, or without AI. For example, the voice generation unit can input scenario content data into a generation AI and have the generation AI perform the application of voice qualities.

[0061] The motion generation unit can estimate the user's emotions and adjust the speed and intensity of the motion based on the estimated emotions. For example, if the user is tense, the motion generation unit can generate slow movements. For example, if the user is relaxed, the motion generation unit can also generate movements at a natural speed. For example, if the user is in a hurry, the motion generation unit can also generate fast and strong movements. This allows for the provision of appropriate motion speed and intensity according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the motion generation unit may be performed using AI, for example, or without AI. For example, the motion generation unit can input user motion data into the generative AI and have the generative AI adjust the speed and intensity of the movements.

[0062] The motion generation unit can apply different motion patterns depending on the scenario content when generating motions. For example, in the case of a severe scenario, the motion generation unit will generate a strict motion. For example, in the case of a gentle scenario, the motion generation unit can also generate a gentle motion. For example, in the case of a neutral scenario, the motion generation unit can also generate a standard motion. This allows for the provision of appropriate motion patterns according to the scenario content. Some or all of the above-described processing in the motion generation unit may be performed using AI, for example, or without AI. For example, the motion generation unit can input scenario content data into a generation AI and have the generation AI execute the application of motion patterns.

[0063] The conversation generation unit can estimate the user's emotions and adjust the content and tone of the conversation based on the estimated emotions. For example, if the user is nervous, the conversation generation unit can generate a concise conversation in a calm tone. For example, if the user is relaxed, the conversation generation unit can also generate a detailed conversation in a bright tone. For example, if the user is in a hurry, the conversation generation unit can also generate a quick and concise conversation. This allows for the provision of appropriate conversation content and tone according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can input user conversation data into the generation AI and have the generation AI adjust the content and tone of the conversation.

[0064] The conversation generation unit can apply different conversation patterns depending on the scenario content when generating conversations. For example, in the case of a severe scenario, the conversation generation unit will generate a strict conversation pattern. For example, in the case of a gentle scenario, the conversation generation unit can also generate a gentle conversation pattern. For example, in the case of a neutral scenario, the conversation generation unit can also generate a standard conversation pattern. This allows for the provision of appropriate conversation patterns according to the content of the scenario. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can input scenario content data into a generation AI and have the generation AI perform the application of conversation patterns.

[0065] The report generation unit can estimate the user's emotions and adjust the report's presentation based on the estimated emotions. For example, if the user is nervous, the report generation unit can generate a concise and to-the-point report. If the user is relaxed, the report generation unit can also generate a report with detailed explanations. If the user is in a hurry, the report generation unit can also generate a quick and concise report. This allows for the provision of an appropriate report presentation style according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processes in the report generation unit may be performed using AI or not. For example, the report generation unit can input user emotion data into the generative AI and have the generative AI adjust the report's presentation style.

[0066] The report generation unit can apply different report formats based on the scenario content and results when generating reports. For example, in the case of a severe scenario, the report generation unit will generate a strict report format. For example, in the case of an easy scenario, the report generation unit can also generate a lenient report format. For example, in the case of a neutral scenario, the report generation unit can also generate a standard report format. This allows for the provision of an appropriate report format according to the scenario content and results. Some or all of the above processing in the report generation unit may be performed using AI, for example, or without AI. For example, the report generation unit can input scenario content data into a generation AI and have the generation AI perform the application of the report format.

[0067] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0068] The customer harassment training system can also include a feedback function. This function provides specific feedback to the user after the simulation is complete. For example, it can explain in detail what went well and what needs improvement in the user's response. Furthermore, the feedback function can estimate the user's emotions and adjust the tone and content of the feedback based on that estimation. For instance, if the user is tense, it can provide feedback in a gentle tone; if the user is relaxed, it can provide detailed feedback. This ensures that appropriate feedback is provided according to the user's emotions.

[0069] The customer harassment training system can also include a progress management unit. This unit manages the user's training progress and suggests the next training session at the appropriate time. For example, it can send a reminder if a user has not received training for a certain period. Furthermore, the progress management unit can estimate the user's emotions and adjust the content and timing of reminders based on these estimations. For instance, it can send a less intense reminder if the user is stressed, and a more proactive reminder if the user is relaxed. This allows for appropriate progress management tailored to the user's emotions.

[0070] The customer harassment training system can also include a rewards section. This section provides rewards to users upon completion of the training. For example, points or badges could be awarded for completing specific scenarios. Furthermore, the rewards section can estimate the user's emotions and adjust the content and timing of rewards based on those estimates. For instance, rewards could be increased if the user is motivated and reduced if they are tired. This allows for the provision of appropriate rewards tailored to the user's emotional state.

[0071] The customer harassment training system can also be equipped with community features. These community features provide a space for users to share their training experiences and opinions. For example, a forum could be created where users can post about what they felt and learned during the training. Furthermore, the community features can estimate users' emotions and filter posts based on those emotions. For instance, if a user has negative emotions, positive posts can be prioritized; if a user has positive emotions, all posts can be displayed. This allows for the provision of an appropriate community environment tailored to the user's emotions.

[0072] The customer harassment training system can also be equipped with a reflection function. This reflection function provides users with a platform to conduct self-assessments after training. For example, a diary function could be included where users can record their feelings and what they learned during the training. Furthermore, the reflection function can estimate the user's emotions and adjust the self-assessment questions based on those estimates. For instance, simple questions could be provided if the user is nervous, while more detailed questions could be provided if the user is relaxed. This allows for an appropriate self-assessment tailored to the user's emotional state.

[0073] The customer harassment training system can also include a real-time assistant. The real-time assistant provides real-time advice to the user during simulations. For example, if the user faces a difficult situation, it can suggest appropriate responses. Furthermore, the real-time assistant can analyze the user's past training history and provide advice based on past successes. For instance, it can suggest responses that worked for the user in the past and instruct the user to avoid responses that failed in the past. This allows for the provision of appropriate advice based on the user's past training history.

[0074] The customer harassment training system can also include a scenario generation unit. This unit automatically generates new scenarios based on the user's job duties and experience. For example, if a user is engaged in a specific task, it can generate scenarios related to that task. Furthermore, the scenario generation unit can analyze the user's past training history and generate variations of previously completed scenarios. For instance, it can adjust the difficulty level of a previously completed scenario and offer it as a new scenario. This allows for the provision of new scenarios tailored to the user's job duties and experience.

[0075] The customer harassment training system can also include a data analysis department. This department analyzes user training data and evaluates the effectiveness of the training. For example, it can compare users' performance before and after training and quantify the training's effectiveness. Furthermore, the data analysis department can predict training effectiveness based on users' past training history. For instance, it can predict the effectiveness of the next training a user should receive based on the effectiveness of their past training. This allows for appropriate evaluation and prediction based on users' training data.

[0076] The customer harassment training system can also include a customization settings section. This section provides users with the ability to freely customize the training settings. For example, users can set the difficulty level of scenarios and the duration of the training. Furthermore, the customization settings section can suggest optimal settings based on the user's past training history. For instance, it can adjust the difficulty level of scenarios the user previously struggled with and increase the difficulty level of scenarios the user excelled at. This allows for the provision of appropriate training settings tailored to the user's needs.

[0077] The customer harassment training system can also be equipped with a multilingual support function. This function allows users to receive training in different languages. For example, users can select scenarios in multiple languages, such as English and Chinese. Furthermore, the multilingual support function can automatically translate scenario content and feedback based on the user's language settings. For example, if a user selects Japanese, all scenarios and feedback will be provided in Japanese. This ensures that appropriate training can be provided to users who speak different languages.

[0078] The following briefly describes the processing flow for example form 2.

[0079] Step 1: The scenario selection section allows the user to select a scenario. The scenario selection section allows the user to choose from multiple scenarios, such as a training scenario or a simulation scenario. Step 2: The voice generation unit generates the customer's voice based on the selected scenario. The voice generation unit generates the customer's voice using, for example, speech synthesis technology. For example, the voice generation unit can adjust the tone and pitch of the customer's voice using speech synthesis AI. Step 3: The motion generation unit generates customer movements based on the selected scenario. The motion generation unit generates full-body movements of the customer, for example, using motion capture technology. For example, the motion generation unit can generate types and patterns of customer movements using video generation AI. Step 4: The conversation generation unit generates customer conversations based on the selected scenario. The conversation generation unit generates customer conversations using, for example, natural language generation technology. For example, the conversation generation unit can use generation AI to generate the content and tone of customer conversations. Step 5: The report generation unit compiles the results generated by the voice generation unit, action generation unit, and conversation generation unit into a report. The report generation unit can, for example, compile the generated results into a text report or a graphical report.

[0080] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0081] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0082] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0083] Each of the multiple elements described above, including the scenario selection unit, voice generation unit, motion generation unit, conversation generation unit, and report generation unit, is implemented in at least one of the smart device 14 and the data processing device 12. For example, the scenario selection unit is implemented by the control unit 46A of the smart device 14, allowing the user to select training scenarios or simulation scenarios. The voice generation unit is implemented by the specific processing unit 290 of the data processing device 12, generating the customer's voice using speech synthesis technology. The motion generation unit is implemented by the control unit 46A of the smart device 14, generating the customer's full-body movements using video generation AI. The conversation generation unit is implemented by the specific processing unit 290 of the data processing device 12, generating the customer's conversation using natural language generation technology. The report generation unit is implemented by the specific processing unit 290 of the data processing device 12, compiling the generated results into a text report or graphical report. The correspondence between each unit and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0084] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0085] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0086] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0087] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0088] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0089] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0090] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0091] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0092] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0093] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0094] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0095] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0096] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0097] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0098] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0099] Each of the multiple elements described above, including the scenario selection unit, voice generation unit, motion generation unit, conversation generation unit, and report generation unit, is implemented in at least one of the smart glasses 214 and the data processing device 12. For example, the scenario selection unit is implemented by the control unit 46A of the smart glasses 214, allowing the user to select training scenarios or simulation scenarios. The voice generation unit is implemented by the specific processing unit 290 of the data processing device 12, for example, and generates the customer's voice using speech synthesis technology. The motion generation unit is implemented by the control unit 46A of the smart glasses 214, for example, and generates the customer's full-body movements using video generation AI. The conversation generation unit is implemented by the specific processing unit 290 of the data processing device 12, for example, and generates the customer's conversation using natural language generation technology. The report generation unit is implemented by the specific processing unit 290 of the data processing device 12, for example, and compiles the generated results as a text report or graphical report. The correspondence between each unit and the device or control unit is not limited to the examples described above, and various changes are possible.

[0100] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0101] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0102] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0103] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0104] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0105] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0106] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0107] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0108] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0109] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0110] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0111] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0112] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0113] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0114] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0115] Each of the multiple elements described above, including the scenario selection unit, voice generation unit, motion generation unit, conversation generation unit, and report generation unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the scenario selection unit is implemented by the control unit 46A of the headset terminal 314, allowing the user to select training scenarios or simulation scenarios. The voice generation unit is implemented by the specific processing unit 290 of the data processing unit 12, generating the customer's voice using speech synthesis technology. The motion generation unit is implemented by the control unit 46A of the headset terminal 314, generating the customer's full-body movements using video generation AI. The conversation generation unit is implemented by the specific processing unit 290 of the data processing unit 12, generating the customer's conversation using natural language generation technology. The report generation unit is implemented by the specific processing unit 290 of the data processing unit 12, compiling the generated results into a text report or graphical report. The correspondence between each unit and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0116] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0117] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0118] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0119] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0120] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0121] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0122] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0123] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0124] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0125] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0126] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0127] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0128] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0129] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0130] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0131] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0132] Each of the multiple elements described above, including the scenario selection unit, voice generation unit, motion generation unit, conversation generation unit, and report generation unit, is implemented in at least one of the following: the robot 414 and the data processing unit 12. For example, the scenario selection unit is implemented by the control unit 46A of the robot 414, allowing the user to select training scenarios or simulation scenarios. The voice generation unit is implemented by the specific processing unit 290 of the data processing unit 12, generating the customer's voice using speech synthesis technology. The motion generation unit is implemented by the control unit 46A of the robot 414, generating the customer's full-body movements using video generation AI. The conversation generation unit is implemented by the specific processing unit 290 of the data processing unit 12, generating the customer's conversation using natural language generation technology. The report generation unit is implemented by the specific processing unit 290 of the data processing unit 12, compiling the generated results into a text report or graphical report. The correspondence between each unit and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0133] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0134] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0135] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0136] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0137] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0138] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0139] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0140] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0141] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0142] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0143] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0144] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0145] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0146] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0147] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0148] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0149] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0150] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0151] (Note 1) A scenario selection section for selecting a scenario, A voice generation unit that generates customer voices based on the scenario selected by the scenario selection unit, An action generation unit that generates customer actions based on the scenario selected by the scenario selection unit, A conversation generation unit that generates customer conversations based on the scenario selected by the scenario selection unit, The system includes a report generation unit that compiles the results generated by the voice generation unit, motion generation unit, and conversation generation unit into a report. A system characterized by the following features. (Note 2) The aforementioned scenario selection unit, The system estimates the user's emotions and adjusts the difficulty of the scenario based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned scenario selection unit, By analyzing past training history, the system automatically suggests the most suitable scenario for the user. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned scenario selection unit, When selecting a scenario, the scenario is customized based on the user's job responsibilities and experience. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned voice generation unit, It estimates the user's emotions and adjusts the tone and speed of the voice based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned voice generation unit, When generating voices, different voice qualities are applied depending on the content of the scenario. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned motion generation unit, It estimates the user's emotions and adjusts the speed and intensity of actions based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned motion generation unit, When generating actions, different action patterns are applied depending on the content of the scenario. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned conversation generation unit, It estimates the user's emotions and adjusts the content and tone of the conversation based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned conversation generation unit, When generating conversations, different conversation patterns are applied depending on the content of the scenario. The system described in Appendix 1, characterized by the features described herein. (Note 11) The report generation unit, It estimates the user's emotions and adjusts how the report is presented based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The report generation unit, When generating reports, apply different report formats based on the scenario content and results. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]

[0152] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A scenario selection section for selecting a scenario, A voice generation unit that generates customer voices based on the scenario selected by the scenario selection unit, An action generation unit that generates customer actions based on the scenario selected by the scenario selection unit, A conversation generation unit that generates customer conversations based on the scenario selected by the scenario selection unit, The system comprises a voice generation unit, an action generation unit, and a report generation unit that compiles the results generated by the speech generation unit, the action generation unit, and the conversation generation unit into a report. A system characterized by the following features.

2. The aforementioned scenario selection unit, The system estimates the user's emotions and adjusts the difficulty of the scenario based on those emotions. The system according to feature 1.

3. The aforementioned scenario selection unit, By analyzing past training history, the system automatically suggests the most suitable scenario for the user. The system according to feature 1.

4. The aforementioned scenario selection unit, When selecting a scenario, the scenario is customized based on the user's job responsibilities and experience. The system according to feature 1.

5. The aforementioned voice generation unit, It estimates the user's emotions and adjusts the tone and speed of the voice based on those emotions. The system according to feature 1.

6. The aforementioned voice generation unit, When generating voices, different voice qualities are applied depending on the content of the scenario. The system according to feature 1.

7. The aforementioned motion generation unit, It estimates the user's emotions and adjusts the speed and intensity of actions based on those emotions. The system according to feature 1.

8. The aforementioned motion generation unit, When generating actions, different action patterns are applied depending on the content of the scenario. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A