system

The telephone call support system addresses poor articulation issues by using AI to read scripts in the user's voice and adjust speed, ensuring clear communication.

JP2026072814APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing communication systems struggle with frequent requests for repetition or paraphrasing due to poor articulation during telephone calls, leading to misunderstandings and miscommunication.

Method used

A telephone call support system that utilizes a generating AI to learn a user's voice, read scripts aloud in their voice, and adjust speaking speed based on listener reactions to enhance clarity and understanding.

Benefits of technology

Enables smoother communication by accurately conveying information in the user's voice and adjusting speaking speed to match listener comprehension, reducing stress and preventing misunderstandings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072814000001_ABST
    Figure 2026072814000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to enable smooth telephone communication even for users with poor articulation. [Solution] The system according to the embodiment comprises a learning unit, a reading unit, and an adjustment unit. The learning unit learns the user's voice. The reading unit reads a script based on the user's voice learned by the learning unit. The adjustment unit adjusts the speaking speed of the voice read by the reading unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0006] , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003] [[ID=​​​​​​​​​​​​​​​​​​​​​​​​​

[0007] The system according to this embodiment enables smooth telephone communication even for users with poor articulation. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The telephone call support system according to an embodiment of the present invention proposes a "read aloud in your own voice option" to solve the problem of frequent requests for repetition or paraphrasing due to poor articulation during telephone calls. The telephone call support system allows a generating AI to learn the user's voice, and the user can have the generating AI read a script at any time during a telephone call. In this case, the generating AI reads the script in the user's voice and adjusts the speaking speed to make it easier for the other party to understand. First, the user's voice is learned by the generating AI. In this process, the user records their own voice and inputs the recording data into the generating AI. The generating AI learns the characteristics of the user's voice and becomes able to generate voices that resemble the user's voice. For example, if the user records "Hello, I'm XX," the generating AI learns that voice and generates voices that resemble the user's voice. Next, the user can have the generating AI read a script at any time during a telephone call. For example, if the user wants to say "The date and time of the next meeting is XX / XX" over the phone, they can prepare that content as a script in advance. During a phone call, if a user instructs the AI ​​to read a script aloud, the AI ​​will read the script in the user's voice. The AI ​​adjusts its speaking speed to ensure clear and easy listening. For example, when the AI ​​reads, "The next meeting is on [date]," it speaks at an appropriate speed to ensure the listener can easily understand. This mechanism solves the problem of frequent clarification and rephrasing due to poor articulation. By having the AI ​​read the script in the user's own voice, users can convey accurate information to the other party. Furthermore, adjusting the speaking speed makes it easier for the listener to understand, facilitating smoother communication. For example, using this function when conveying important information in a business setting can prevent misunderstandings and miscommunication. This function is also offered as a call option by communication service providers. Users can utilize the AI-powered reading function during phone calls by using these services. For example, users of communication service providers can use this option to have their scripts read aloud in their own voice during phone calls. This reduces the stress of phone calls and enables more comfortable communication.This allows the telephone call support system to learn the user's voice, read it aloud, and adjust the speaking speed, thereby solving the problem of slurred speech during phone calls.

[0029] The telephone call support system according to this embodiment comprises a learning unit, a reading unit, and an adjustment unit. The learning unit learns the user's voice. For example, the learning unit learns the characteristics of the user's voice by having the user record their voice and input the recording data into the generation AI. The generation AI learns the characteristics of the user's voice, such as tone, pitch, and intonation, and can generate speech that resembles the user's voice. For example, if the user records "Hello, I'm XX," the generation AI learns that voice and generates speech that resembles the user's voice. The reading unit reads a script based on the user's voice learned by the learning unit. For example, if the user wants to say "The date and time of the next meeting is XX / XX" over the phone, the reading unit prepares the content as a script in advance. During a telephone call, if the user instructs the generation AI to read the script, the generation AI reads the script in the user's voice. The generation AI can generate natural speech while retaining the characteristics of the user's voice. The adjustment unit adjusts the speaking speed of the speech read by the reading unit. The adjustment unit adjusts the speaking speed based, for example, on the other party's reaction speed and speech clarity. The adjustment unit uses a generation AI to speak at an appropriate speed so that the other party can easily understand it. For example, when the generation AI reads aloud, "The date and time of the next meeting is [date]", it speaks at an appropriate speed so that the other party can easily understand it. In this way, the telephone call support system according to the embodiment can solve the problem of slurred speech during telephone calls by learning the user's voice, reading aloud, and adjusting the speaking speed.

[0030] The learning unit learns the user's voice. For example, the learning unit learns the characteristics of the user's voice by inputting the recording data into the generating AI. Specifically, the user records their voice using a smartphone or a dedicated recording device. This recording data is uploaded to a cloud server so that the generating AI can access it. The generating AI analyzes the audio data and extracts features such as the tone, pitch, intonation, and pronunciation habits of the user's voice. These features are stored as a voice model and used for subsequent voice generation. For example, if a user records "Hello, I'm XX," the generating AI analyzes that audio data and learns the characteristics of the user's voice. Using deep learning technology, the generating AI can model the characteristics of the user's voice with high accuracy and generate voices that resemble the user's voice. Furthermore, the learning unit can collect multiple audio data of the user speaking in different situations and with different emotions, learning a wider variety of voice characteristics. As a result, the generating AI has a rich variety of user voices, enabling natural-sounding voice generation. The learning unit can adapt to changes in the user's voice by regularly collecting new voice data from the user and updating the model. This ensures that the learning unit always retains the latest characteristics of the user's voice, enabling highly accurate speech generation.

[0031] The reading unit reads aloud a script based on the user's voice, which has been learned by the learning unit. For example, if a user wants to say "The date and time of the next meeting is [date]" over the phone, the reading unit prepares the content as a script in advance. The user uses a smartphone or computer to input the script in text format and saves it to the system. During a phone call, the user instructs the generating AI to read the script, and the generating AI reads the script in the user's voice. The generating AI can generate natural-sounding speech while preserving the characteristics of the user's voice. Specifically, when converting text data into audio data, the generating AI applies the tone, pitch, and intonation of the user's voice to produce speech that sounds as if the user is actually speaking. Furthermore, the reading unit can add appropriate emotional expressions and emphasis depending on the content of the script. For example, when conveying important information, it can emphasize the tone of voice to make it clear to the listener. The reading unit utilizes the speech generation capabilities of the generating AI to achieve both naturalness and clarity of the user's voice. This allows users to smoothly convey information in their own voice during phone calls, improving the efficiency of communication.

[0032] The adjustment unit adjusts the speaking speed of the audio read aloud by the reading unit. For example, the adjustment unit adjusts the speaking speed based on the other party's reaction speed and speech clarity. Specifically, the adjustment unit monitors the other party's reactions in real time and automatically adjusts the speaking speed to make the audio easier for them to understand. The generation AI analyzes the audio data and evaluates the other party's reaction time and speech clarity. For example, if the other party does not immediately react to the audio reading "The date and time of the next meeting is [date]", the adjustment unit slows down the speaking speed to make it easier for the other party to understand. Also, if the other party cannot hear clearly, the adjustment unit adjusts the pitch and intonation of the audio to improve clarity. Furthermore, the adjustment unit can also customize the speaking speed according to the user's preferences. Users can adjust the speaking speed and tone of voice from the system settings screen and select the optimal audio settings for themselves. This allows the adjustment unit to provide an optimal audio environment for both the user and the other party, enabling smooth communication. The adjustment unit utilizes advanced voice analysis technology from generation AI to adjust the voice in real time, ensuring that the voice quality during a call is always optimal.

[0033] The recording unit can record the user's voice. For example, the recording unit can provide data for the generating AI to learn the characteristics of the user's voice by having the user record their own voice and inputting that recording data into the AI. The recording unit can select the optimal recording settings depending on the type of microphone used and the recording environment. For example, the recording unit can record clear audio using a high-quality microphone. In addition, the recording unit can generate low-noise audio data by recording in a quiet environment. In this way, the recording unit can provide data for the generating AI to learn from by recording the user's voice.

[0034] The input unit can input data recorded by the recording unit into the generating AI. For example, the input unit can teach the generating AI the characteristics of the user's voice by inputting the recorded data. The input unit can select the optimal input method depending on the data format and timing of input. For example, the input unit can input the recorded data into the generating AI in text or voice format. Furthermore, the input unit can achieve rapid learning by inputting the recorded data into the generating AI in real time. As a result, the input unit can teach the generating AI the characteristics of the user's voice by inputting the recorded data.

[0035] The adjustment unit may include a reference unit that adjusts the speaking speed based on the listener's reaction speed and speech clarity. For example, the adjustment unit monitors the listener's reaction speed in real time and maintains an optimal speaking speed. The adjustment unit uses generative AI to speak at an appropriate speed so that the listener can easily understand it. For example, if the listener reacts slowly, the speaking speed will be slowed down. Conversely, if the listener reacts quickly, the speaking speed can be increased. Furthermore, the adjustment unit can automatically evaluate speech clarity and readjust the speed as needed. For example, if the speech is unclear, the speaking speed will be slowed down to improve clarity. In this way, the adjustment unit can improve intelligibility by adjusting the speaking speed based on the listener's reaction speed and speech clarity.

[0036] The learning unit can analyze the user's voice tone and intonation in detail during learning to generate more natural-sounding speech. For example, the learning unit can analyze the pitch of the user's voice in detail to reproduce natural intonation. It can also analyze the speed and rhythm of the user's voice to reproduce natural speaking. Furthermore, it can analyze the volume of the user's voice to enrich emotional expression. In this way, the learning unit can generate natural-sounding speech by analyzing the user's voice tone and intonation in detail.

[0037] The learning unit can track changes in the user's voice in real time during learning and reflect the latest voice characteristics. For example, if the user catches a cold, the learning unit will learn the changes in their voice in real time. It can also learn the changes in the user's voice in real time if they are nervous. Furthermore, it can learn the changes in the user's voice in real time if they are relaxed. In this way, the learning unit can reflect the latest voice characteristics by tracking changes in the user's voice in real time.

[0038] The learning unit can remove background noise from the user's voice during training, generating clear audio data. For example, if the user records in a noisy environment, the learning unit can remove background noise to produce clear audio. It can also remove wind noise if the user records in a windy location, and further remove echo if the user records indoors. In this way, the learning unit can generate clear audio data by removing background noise.

[0039] The learning unit can learn different languages ​​and dialects from the user's voice during training, enabling multilingual support. For example, if the user speaks English and Japanese, the learning unit can learn both languages ​​to achieve multilingual support. Furthermore, if the user speaks Kansai dialect, it can learn that dialect to generate natural-sounding speech. Additionally, if the user speaks French, it can learn that language to achieve multilingual support. In this way, the learning unit can achieve multilingual support by learning different languages ​​and dialects.

[0040] The narrator can add appropriate intonation and emphasis during reading, depending on the content of the script. For example, the narrator can emphasize important information, use intonation when conveying emotional content, and read in a calm tone when conveying relaxed content. In this way, the narrator can improve intelligibility by adding intonation and emphasis according to the content of the script.

[0041] The narration unit can select different voice styles while maintaining the characteristics of the user's voice. For example, in a business setting, it can narrate in a formal voice style. In a casual conversation, it can narrate in a relaxed voice style. Furthermore, in a presentation, it can narrate in a clear and powerful voice style. In this way, the narration unit can adapt to various situations by selecting different voice styles while maintaining the characteristics of the user's voice.

[0042] The narrator can improve intelligibility by inserting appropriate pauses based on the content of the script during reading. For example, the narrator can insert appropriate pauses before conveying important information. It can also enhance the effect of conveying emotional content by inserting pauses. Furthermore, it can improve intelligibility when reading long texts by inserting appropriate pauses. In this way, the narrator can improve intelligibility by inserting appropriate pauses based on the content of the script.

[0043] The narration function can enhance the quality of the voice by adding echo and reverb effects to the user's voice during narration. For example, it can enhance the quality of voice by adding echo effects when making an important presentation. It can also enhance the quality of voice by adding reverb effects when conveying emotional content. Furthermore, it can enhance the quality of voice by adding echo effects when giving a presentation. In short, the narration function can improve the quality of voice by adding echo and reverb effects.

[0044] The adjustment unit can monitor the other party's reaction speed in real time during adjustment and maintain an optimal speaking speed. For example, if the other party reacts slowly, the adjustment unit will slow down the speaking speed. Conversely, if the other party reacts quickly, it can also speed up the speaking speed. Furthermore, it can adjust the speaking speed in real time according to the other party's reaction speed. In this way, the adjustment unit can maintain an optimal speaking speed by monitoring the other party's reaction speed in real time.

[0045] The adjustment unit automatically evaluates the clarity of the speech during adjustment and can readjust the speed as needed. For example, if the speech is unclear, the adjustment unit can slow down the speaking speed to improve clarity. Conversely, if the speech is clear, it can also increase the speaking speed to improve efficiency. Furthermore, it can automatically readjust the speaking speed according to the clarity of the speech. In this way, the adjustment unit can automatically evaluate the clarity of the speech and readjust the speed as needed.

[0046] The adjustment unit can customize the speaking speed based on the listener's hearing characteristics during adjustment. For example, if the listener is elderly, the adjustment unit will speak at a slower pace. If the listener is young, it can speak at a natural pace. Furthermore, it can customize the speaking speed according to the listener's hearing characteristics. In this way, the adjustment unit can achieve more appropriate communication by customizing the speaking speed based on the listener's hearing characteristics.

[0047] The adjustment unit can detect the level of background noise during adjustment and apply noise cancellation to improve speech clarity. For example, when speaking in a noisy environment, the adjustment unit can apply noise cancellation to improve speech clarity. Conversely, when speaking in a quiet environment, it can maintain natural speech without applying noise cancellation. Furthermore, it can automatically apply noise cancellation according to the level of background noise. In this way, the adjustment unit can detect the level of background noise and apply noise cancellation to improve speech clarity.

[0048] The recording unit can automatically evaluate the quality of the user's voice during recording and select the optimal recording settings. For example, if the user's voice is not clear, the recording unit will apply noise reduction. It can also automatically adjust the volume if the user's voice is quiet. Furthermore, if the user's voice is echoing, it can apply echo cancellation. In this way, the recording unit can automatically evaluate the quality of the user's voice and select the optimal recording settings.

[0049] The recording unit can use multiple microphones to record spatial audio during recording, thereby generating more natural-sounding audio data. For example, the recording unit can use multiple microphones to record the user's voice spatially. It can also use multiple microphones to record background sounds spatially. Furthermore, it can use multiple microphones to record ambient sounds spatially. As a result, the recording unit can generate more natural-sounding audio data by using multiple microphones to record spatial audio.

[0050] The input unit can analyze the characteristics of the user's voice in real time during input and select the optimal input method. For example, if the user's voice is unclear, the input unit will avoid voice input. Conversely, if the user's voice is clear, it can prioritize voice input. Furthermore, it can select the optimal input method based on the characteristics of the user's voice. In this way, the input unit can select the optimal input method by analyzing the characteristics of the user's voice in real time.

[0051] The input unit can remove background noise from the user's voice during input, generating clear audio data. For example, if the user is inputting in a noisy environment, the input unit can remove background noise. It can also remove wind noise if the user is inputting in a windy location. Furthermore, it can remove echoes if the user is inputting indoors. In this way, the input unit can generate clear audio data by removing background noise.

[0052] The reference unit can maintain optimal standards by monitoring the other party's reaction speed and voice clarity in real time when setting the standards. For example, if the other party's reaction speed is slow, the reference unit can relax the standards. Conversely, if the other party's reaction speed is fast, it can tighten the standards. Furthermore, it can adjust the standards according to the other party's voice clarity. In this way, the reference unit can maintain optimal standards by monitoring the other party's reaction speed and voice clarity in real time.

[0053] The reference unit can customize the criteria based on the other party's hearing characteristics when setting the criteria. For example, if the other party is elderly, the reference unit will set criteria that are appropriate to their hearing characteristics. If the other party is young, it can also set natural criteria. Furthermore, it can customize the criteria according to the other party's hearing characteristics. In this way, the reference unit can set more appropriate criteria by customizing the criteria based on the other party's hearing characteristics.

[0054] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0055] The telephone call support system can translate conversations in real time, facilitating communication between users who speak different languages. For example, when a Japanese-speaking user and an English-speaking user are on a call, the conversation is translated in real time and conveyed to the other party. Similarly, translation can be performed when a French-speaking user and a Spanish-speaking user are on a call. Furthermore, even when users who speak different dialects are on a call, the dialect can be converted into standard Japanese. This facilitates smooth communication between users who speak different languages ​​or dialects.

[0056] A telephone call support system can automatically record what a user says during a call, allowing for later review. For example, it can record important statements during a business meeting for later review. It can also record important information during calls with family and friends. Furthermore, even if it's difficult to take notes during a call, the automatic recording function allows for later review. This ensures that users don't miss important information during calls and can review it later.

[0057] The telephone call support system can transcribe the user's speech in real time during a call, allowing for visual confirmation. For example, if a user with a hearing impairment is making a call, their speech will be transcribed and displayed on the screen. Similarly, even in noisy environments, the system can transcribe speech for confirmation. Furthermore, even if taking notes during a call is difficult, the real-time transcription function allows for visual confirmation of what is being said. This enables users to visually confirm what they are saying during a call, leading to more effective communication.

[0058] A telephone call support system can analyze a user's statements during a call and suggest appropriate responses. For example, it can suggest appropriate answers during a business meeting, supporting smooth progress. It can also suggest appropriate responses during customer support calls. Furthermore, it can suggest appropriate topics during calls with friends and family. This allows users to receive appropriate responses during calls, leading to more effective communication.

[0059] A telephone call support system can analyze user statements during a call and evaluate the call's progress in real time. For example, it can evaluate the progress of an agenda item during a business meeting and suggest the next steps. It can also evaluate the progress of problem-solving during a customer support call and suggest the next course of action. Furthermore, it can evaluate the progress of a conversation during a call with friends or family and suggest the next topic. This allows users to understand the progress of their calls in real time, enabling more effective communication.

[0060] The following briefly describes the processing flow for example form 1.

[0061] Step 1: The learning unit learns the user's voice. The user records their voice, and this recording data is input into the generating AI, which learns the characteristics of the user's voice. The generating AI learns features such as the tone, pitch, and intonation of the user's voice, and can generate speech that resembles the user's voice. Step 2: The reading unit reads the script based on the user's voice, which has been learned by the learning unit. When the user instructs the generating AI to read a script they have prepared in advance, the generating AI reads the script in the user's voice. The generating AI can produce natural-sounding speech while retaining the characteristics of the user's voice. Step 3: The adjustment unit adjusts the speaking speed of the audio read aloud by the reading unit. It adjusts the speaking speed based on the listener's reaction speed and speech clarity, and uses a generation AI to speak at an appropriate speed that is easy for the listener to understand.

[0062] (Example of form 2) The telephone call support system according to an embodiment of the present invention proposes a "read aloud in your own voice option" to solve the problem of frequent requests for repetition or paraphrasing due to poor articulation during telephone calls. The telephone call support system allows a generating AI to learn the user's voice, and the user can have the generating AI read a script at any time during a telephone call. In this case, the generating AI reads the script in the user's voice and adjusts the speaking speed to make it easier for the other party to understand. First, the user's voice is learned by the generating AI. In this process, the user records their own voice and inputs the recording data into the generating AI. The generating AI learns the characteristics of the user's voice and becomes able to generate voices that resemble the user's voice. For example, if the user records "Hello, I'm XX," the generating AI learns that voice and generates voices that resemble the user's voice. Next, the user can have the generating AI read a script at any time during a telephone call. For example, if the user wants to say "The date and time of the next meeting is XX / XX" over the phone, they can prepare that content as a script in advance. During a phone call, if a user instructs the AI ​​to read a script aloud, the AI ​​will read the script in the user's voice. The AI ​​adjusts its speaking speed to ensure clear and easy listening. For example, when the AI ​​reads, "The next meeting is on [date]," it speaks at an appropriate speed to ensure the listener can easily understand. This mechanism solves the problem of frequent clarification and rephrasing due to poor articulation. By having the AI ​​read the script in the user's own voice, users can convey accurate information to the other party. Furthermore, adjusting the speaking speed makes it easier for the listener to understand, facilitating smoother communication. For example, using this function when conveying important information in a business setting can prevent misunderstandings and miscommunication. This function is also offered as a call option by communication service providers. Users can utilize the AI-powered reading function during phone calls by using these services. For example, users of communication service providers can use this option to have their scripts read aloud in their own voice during phone calls. This reduces the stress of phone calls and enables more comfortable communication.This allows the telephone call support system to learn the user's voice, read it aloud, and adjust the speaking speed, thereby solving the problem of slurred speech during phone calls.

[0063] The telephone call support system according to this embodiment comprises a learning unit, a reading unit, and an adjustment unit. The learning unit learns the user's voice. For example, the learning unit learns the characteristics of the user's voice by having the user record their voice and input the recording data into the generation AI. The generation AI learns the characteristics of the user's voice, such as tone, pitch, and intonation, and can generate speech that resembles the user's voice. For example, if the user records "Hello, I'm XX," the generation AI learns that voice and generates speech that resembles the user's voice. The reading unit reads a script based on the user's voice learned by the learning unit. For example, if the user wants to say "The date and time of the next meeting is XX / XX" over the phone, the reading unit prepares the content as a script in advance. During a telephone call, if the user instructs the generation AI to read the script, the generation AI reads the script in the user's voice. The generation AI can generate natural speech while retaining the characteristics of the user's voice. The adjustment unit adjusts the speaking speed of the speech read by the reading unit. The adjustment unit adjusts the speaking speed based, for example, on the other party's reaction speed and speech clarity. The adjustment unit uses a generation AI to speak at an appropriate speed so that the other party can easily understand it. For example, when the generation AI reads aloud, "The date and time of the next meeting is [date]", it speaks at an appropriate speed so that the other party can easily understand it. In this way, the telephone call support system according to the embodiment can solve the problem of slurred speech during telephone calls by learning the user's voice, reading aloud, and adjusting the speaking speed.

[0064] The learning unit learns the user's voice. For example, the learning unit learns the characteristics of the user's voice by inputting the recording data into the generating AI. Specifically, the user records their voice using a smartphone or a dedicated recording device. This recording data is uploaded to a cloud server so that the generating AI can access it. The generating AI analyzes the audio data and extracts features such as the tone, pitch, intonation, and pronunciation habits of the user's voice. These features are stored as a voice model and used for subsequent voice generation. For example, if a user records "Hello, I'm XX," the generating AI analyzes that audio data and learns the characteristics of the user's voice. Using deep learning technology, the generating AI can model the characteristics of the user's voice with high accuracy and generate voices that resemble the user's voice. Furthermore, the learning unit can collect multiple audio data of the user speaking in different situations and with different emotions, learning a wider variety of voice characteristics. As a result, the generating AI has a rich variety of user voices, enabling natural-sounding voice generation. The learning unit can adapt to changes in the user's voice by regularly collecting new voice data from the user and updating the model. This ensures that the learning unit always retains the latest characteristics of the user's voice, enabling highly accurate speech generation.

[0065] The reading unit reads aloud a script based on the user's voice, which has been learned by the learning unit. For example, if a user wants to say "The date and time of the next meeting is [date]" over the phone, the reading unit prepares the content as a script in advance. The user uses a smartphone or computer to input the script in text format and saves it to the system. During a phone call, the user instructs the generating AI to read the script, and the generating AI reads the script in the user's voice. The generating AI can generate natural-sounding speech while preserving the characteristics of the user's voice. Specifically, when converting text data into audio data, the generating AI applies the tone, pitch, and intonation of the user's voice to produce speech that sounds as if the user is actually speaking. Furthermore, the reading unit can add appropriate emotional expressions and emphasis depending on the content of the script. For example, when conveying important information, it can emphasize the tone of voice to make it clear to the listener. The reading unit utilizes the speech generation capabilities of the generating AI to achieve both naturalness and clarity of the user's voice. This allows users to smoothly convey information in their own voice during phone calls, improving the efficiency of communication.

[0066] The adjustment unit adjusts the speaking speed of the audio read aloud by the reading unit. For example, the adjustment unit adjusts the speaking speed based on the other party's reaction speed and speech clarity. Specifically, the adjustment unit monitors the other party's reactions in real time and automatically adjusts the speaking speed to make the audio easier for them to understand. The generation AI analyzes the audio data and evaluates the other party's reaction time and speech clarity. For example, if the other party does not immediately react to the audio reading "The date and time of the next meeting is [date]", the adjustment unit slows down the speaking speed to make it easier for the other party to understand. Also, if the other party cannot hear clearly, the adjustment unit adjusts the pitch and intonation of the audio to improve clarity. Furthermore, the adjustment unit can also customize the speaking speed according to the user's preferences. Users can adjust the speaking speed and tone of voice from the system settings screen and select the optimal audio settings for themselves. This allows the adjustment unit to provide an optimal audio environment for both the user and the other party, enabling smooth communication. The adjustment unit utilizes advanced voice analysis technology from generation AI to adjust the voice in real time, ensuring that the voice quality during a call is always optimal.

[0067] The recording unit can record the user's voice. For example, the recording unit can provide data for the generating AI to learn the characteristics of the user's voice by having the user record their own voice and inputting that recording data into the AI. The recording unit can select the optimal recording settings depending on the type of microphone used and the recording environment. For example, the recording unit can record clear audio using a high-quality microphone. In addition, the recording unit can generate low-noise audio data by recording in a quiet environment. In this way, the recording unit can provide data for the generating AI to learn from by recording the user's voice.

[0068] The input unit can input data recorded by the recording unit into the generating AI. For example, the input unit can teach the generating AI the characteristics of the user's voice by inputting the recorded data. The input unit can select the optimal input method depending on the data format and timing of input. For example, the input unit can input the recorded data into the generating AI in text or voice format. Furthermore, the input unit can achieve rapid learning by inputting the recorded data into the generating AI in real time. As a result, the input unit can teach the generating AI the characteristics of the user's voice by inputting the recorded data.

[0069] The adjustment unit may include a reference unit that adjusts the speaking speed based on the listener's reaction speed and speech clarity. For example, the adjustment unit monitors the listener's reaction speed in real time and maintains an optimal speaking speed. The adjustment unit uses generative AI to speak at an appropriate speed so that the listener can easily understand it. For example, if the listener reacts slowly, the speaking speed will be slowed down. Conversely, if the listener reacts quickly, the speaking speed can be increased. Furthermore, the adjustment unit can automatically evaluate speech clarity and readjust the speed as needed. For example, if the speech is unclear, the speaking speed will be slowed down to improve clarity. In this way, the adjustment unit can improve intelligibility by adjusting the speaking speed based on the listener's reaction speed and speech clarity.

[0070] The learning unit can estimate the user's emotions and select training data based on the estimated emotions. For example, if the user is nervous, the learning unit will prioritize training on audio samples that promote relaxation. If the user is relaxed, it can also train on audio samples of natural conversation. Furthermore, if the user is in a hurry, it can select audio samples that can be learned in a short amount of time. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows the learning unit to train on more appropriate audio samples by selecting training data based on the user's emotions.

[0071] The learning unit can analyze the user's voice tone and intonation in detail during learning to generate more natural-sounding speech. For example, the learning unit can analyze the pitch of the user's voice in detail to reproduce natural intonation. It can also analyze the speed and rhythm of the user's voice to reproduce natural speaking. Furthermore, it can analyze the volume of the user's voice to enrich emotional expression. In this way, the learning unit can generate natural-sounding speech by analyzing the user's voice tone and intonation in detail.

[0072] The learning unit can track changes in the user's voice in real time during learning and reflect the latest voice characteristics. For example, if the user catches a cold, the learning unit will learn the changes in their voice in real time. It can also learn the changes in the user's voice in real time if they are nervous. Furthermore, it can learn the changes in the user's voice in real time if they are relaxed. In this way, the learning unit can reflect the latest voice characteristics by tracking changes in the user's voice in real time.

[0073] The learning unit can estimate the user's emotions and adjust the learning frequency based on the estimated emotions. For example, if the user is stressed, the learning unit can reduce the learning frequency to alleviate the burden. Conversely, if the user is relaxed, it can increase the learning frequency to learn more efficiently. Furthermore, if the user is in a hurry, it can adjust the frequency to complete the learning in a shorter time. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows the learning unit to learn efficiently while reducing the burden by adjusting the learning frequency based on the user's emotions.

[0074] The learning unit can remove background noise from the user's voice during training, generating clear audio data. For example, if the user records in a noisy environment, the learning unit can remove background noise to produce clear audio. It can also remove wind noise if the user records in a windy location, and further remove echo if the user records indoors. In this way, the learning unit can generate clear audio data by removing background noise.

[0075] The learning unit can learn different languages ​​and dialects from the user's voice during training, enabling multilingual support. For example, if the user speaks English and Japanese, the learning unit can learn both languages ​​to achieve multilingual support. Furthermore, if the user speaks Kansai dialect, it can learn that dialect to generate natural-sounding speech. Additionally, if the user speaks French, it can learn that language to achieve multilingual support. In this way, the learning unit can achieve multilingual support by learning different languages ​​and dialects.

[0076] The narration unit can estimate the user's emotions and adjust the tone and emotional expression of the narration based on the estimated emotions. For example, if the user is nervous, the narration unit will narrate in a calm tone. If the user is relaxed, it can narrate in a bright tone. Furthermore, if the user is in a hurry, it can narrate in a quick and concise tone. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows the narration unit to achieve more natural narration by adjusting the tone and emotional expression of the narration based on the user's emotions.

[0077] The narrator can add appropriate intonation and emphasis during reading, depending on the content of the script. For example, the narrator can emphasize important information, use intonation when conveying emotional content, and read in a calm tone when conveying relaxed content. In this way, the narrator can improve intelligibility by adding intonation and emphasis according to the content of the script.

[0078] The narration unit can select different voice styles while maintaining the characteristics of the user's voice. For example, in a business setting, it can narrate in a formal voice style. In a casual conversation, it can narrate in a relaxed voice style. Furthermore, in a presentation, it can narrate in a clear and powerful voice style. In this way, the narration unit can adapt to various situations by selecting different voice styles while maintaining the characteristics of the user's voice.

[0079] The narration unit can estimate the user's emotions and adjust the reading speed based on the estimated emotions. For example, if the user is nervous, the narration unit will read slowly. If the user is relaxed, it can read at a natural speed. Furthermore, if the user is in a hurry, it can read at a rapid speed. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows the narration unit to read at a more appropriate speed by adjusting the reading speed based on the user's emotions.

[0080] The narrator can improve intelligibility by inserting appropriate pauses based on the content of the script during reading. For example, the narrator can insert appropriate pauses before conveying important information. It can also enhance the effect of conveying emotional content by inserting pauses. Furthermore, it can improve intelligibility when reading long texts by inserting appropriate pauses. In this way, the narrator can improve intelligibility by inserting appropriate pauses based on the content of the script.

[0081] The narration function can enhance the quality of the voice by adding echo and reverb effects to the user's voice during narration. For example, it can enhance the quality of voice by adding echo effects when making an important presentation. It can also enhance the quality of voice by adding reverb effects when conveying emotional content. Furthermore, it can enhance the quality of voice by adding echo effects when giving a presentation. In short, the narration function can improve the quality of voice by adding echo and reverb effects.

[0082] The adjustment unit can estimate the user's emotions and adjust the speaking speed based on the estimated emotions. For example, if the user is nervous, the adjustment unit will speak slowly. If the user is relaxed, it can speak at a natural speed. Furthermore, if the user is in a hurry, it can speak at a rapid speed. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows the adjustment unit to speak at a more appropriate speed by adjusting the speaking speed based on the user's emotions.

[0083] The adjustment unit can monitor the other party's reaction speed in real time during adjustment and maintain an optimal speaking speed. For example, if the other party reacts slowly, the adjustment unit will slow down the speaking speed. Conversely, if the other party reacts quickly, it can also speed up the speaking speed. Furthermore, it can adjust the speaking speed in real time according to the other party's reaction speed. In this way, the adjustment unit can maintain an optimal speaking speed by monitoring the other party's reaction speed in real time.

[0084] The adjustment unit automatically evaluates the clarity of the speech during adjustment and can readjust the speed as needed. For example, if the speech is unclear, the adjustment unit can slow down the speaking speed to improve clarity. Conversely, if the speech is clear, it can also increase the speaking speed to improve efficiency. Furthermore, it can automatically readjust the speaking speed according to the clarity of the speech. In this way, the adjustment unit can automatically evaluate the clarity of the speech and readjust the speed as needed.

[0085] The adjustment unit can estimate the user's emotions and adjust the volume based on those emotions. For example, if the user is nervous, the adjustment unit can lower the volume to create a calm atmosphere. Conversely, if the user is relaxed, it can raise the volume to create a cheerful atmosphere. Furthermore, if the user is in a hurry, it can adjust the volume to facilitate quick communication. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows the adjustment unit to speak at a more appropriate volume by adjusting the volume based on the user's emotions.

[0086] The adjustment unit can customize the speaking speed based on the listener's hearing characteristics during adjustment. For example, if the listener is elderly, the adjustment unit will speak at a slower pace. If the listener is young, it can speak at a natural pace. Furthermore, it can customize the speaking speed according to the listener's hearing characteristics. In this way, the adjustment unit can achieve more appropriate communication by customizing the speaking speed based on the listener's hearing characteristics.

[0087] The adjustment unit can detect the level of background noise during adjustment and apply noise cancellation to improve speech clarity. For example, when speaking in a noisy environment, the adjustment unit can apply noise cancellation to improve speech clarity. Conversely, when speaking in a quiet environment, it can maintain natural speech without applying noise cancellation. Furthermore, it can automatically apply noise cancellation according to the level of background noise. In this way, the adjustment unit can detect the level of background noise and apply noise cancellation to improve speech clarity.

[0088] The recording unit can estimate the user's emotions and adjust the recording start time based on the estimated emotions. For example, if the user is tense, the recording unit will not start recording until the user is relaxed. Conversely, if the user is relaxed, it can start recording immediately. Furthermore, if the user is in a hurry, it can start recording quickly. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows the recording unit to start recording at a more appropriate time by adjusting the recording start time based on the user's emotions.

[0089] The recording unit can automatically evaluate the quality of the user's voice during recording and select the optimal recording settings. For example, if the user's voice is not clear, the recording unit will apply noise reduction. It can also automatically adjust the volume if the user's voice is quiet. Furthermore, if the user's voice is echoing, it can apply echo cancellation. In this way, the recording unit can automatically evaluate the quality of the user's voice and select the optimal recording settings.

[0090] The recording unit can estimate the user's emotions and adjust the recording length based on the estimated emotions. For example, if the user is nervous, the recording unit may recommend a shorter recording. Conversely, if the user is relaxed, it may recommend a longer recording. Furthermore, if the user is in a hurry, it can complete the recording quickly. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows the recording unit to record at a more appropriate length by adjusting the recording length based on the user's emotions.

[0091] The recording unit can use multiple microphones to record spatial audio during recording, thereby generating more natural-sounding audio data. For example, the recording unit can use multiple microphones to record the user's voice spatially. It can also use multiple microphones to record background sounds spatially. Furthermore, it can use multiple microphones to record ambient sounds spatially. As a result, the recording unit can generate more natural-sounding audio data by using multiple microphones to record spatial audio.

[0092] The input unit can estimate the user's emotions and prioritize input data based on those emotions. For example, if the user is stressed, the input unit will prioritize inputting important data. If the user is relaxed, it can also prioritize inputting detailed data. Furthermore, if the user is in a hurry, it can prioritize data that can be entered quickly. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows the input unit to prioritize inputting more important data by prioritizing input data based on the user's emotions.

[0093] The input unit can analyze the characteristics of the user's voice in real time during input and select the optimal input method. For example, if the user's voice is unclear, the input unit will avoid voice input. Conversely, if the user's voice is clear, it can prioritize voice input. Furthermore, it can select the optimal input method based on the characteristics of the user's voice. In this way, the input unit can select the optimal input method by analyzing the characteristics of the user's voice in real time.

[0094] The input unit can estimate the user's emotions and filter the input data based on those emotions. For example, if the user is tense, the input unit can filter only the important data. If the user is relaxed, it can filter for more detailed data. Furthermore, if the user is in a hurry, it can filter for data that can be entered quickly. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows the input unit to prioritize processing of more important data by filtering the input data based on the user's emotions.

[0095] The input unit can remove background noise from the user's voice during input, generating clear audio data. For example, if the user is inputting in a noisy environment, the input unit can remove background noise. It can also remove wind noise if the user is inputting in a windy location. Furthermore, it can remove echoes if the user is inputting indoors. In this way, the input unit can generate clear audio data by removing background noise.

[0096] The criteria unit can estimate the user's emotions and adjust the criteria settings based on the estimated emotions. For example, if the user is tense, the criteria unit can set a relaxing criterion. If the user is relaxed, it can also set a natural criterion. Furthermore, if the user is in a hurry, it can set a criterion that allows for a quick response. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows the criteria unit to set more appropriate criteria by adjusting the criteria settings based on the user's emotions.

[0097] The reference unit can maintain optimal standards by monitoring the other party's reaction speed and voice clarity in real time when setting the standards. For example, if the other party's reaction speed is slow, the reference unit can relax the standards. Conversely, if the other party's reaction speed is fast, it can tighten the standards. Furthermore, it can adjust the standards according to the other party's voice clarity. In this way, the reference unit can maintain optimal standards by monitoring the other party's reaction speed and voice clarity in real time.

[0098] The criteria unit can estimate the user's emotions and determine the priority of criteria based on the estimated emotions. For example, if the user is stressed, the criteria unit will prioritize important criteria. If the user is relaxed, it may also prioritize detailed criteria. Furthermore, if the user is in a hurry, it may prioritize criteria that allow for a quick response. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows the criteria unit to prioritize more important criteria by determining the priority of criteria based on the user's emotions.

[0099] The reference unit can customize the criteria based on the other party's hearing characteristics when setting the criteria. For example, if the other party is elderly, the reference unit will set criteria that are appropriate to their hearing characteristics. If the other party is young, it can also set natural criteria. Furthermore, it can customize the criteria according to the other party's hearing characteristics. In this way, the reference unit can set more appropriate criteria by customizing the criteria based on the other party's hearing characteristics.

[0100] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0101] A telephone call support system can estimate a user's emotions and provide a summary of the call based on those emotions. For example, if the user is nervous, the summary will be concise and clear. If the user is relaxed, a detailed summary can be provided. Furthermore, if the user is in a hurry, a summary can be provided quickly. This allows users to receive the most appropriate summary according to their emotions and efficiently understand the content of the call.

[0102] A telephone call support system can analyze the user's voice tone and intonation and provide appropriate feedback during a call. For example, if the user is nervous, it can offer advice to help them relax. If the user is relaxed, it can also provide feedback to help them continue the conversation naturally. Furthermore, if the user is in a hurry, it can offer advice to help them speak more quickly. This allows users to receive appropriate feedback during calls and achieve more effective communication.

[0103] A telephone call support system can estimate the user's emotions and suggest an appropriate time to end the call based on those emotions. For example, if the user is feeling nervous, it can suggest ending the call early. If the user is relaxed, it can suggest continuing the call. Furthermore, if the user is in a hurry, it can suggest ending the call quickly. This allows users to end the call at the optimal time according to their emotions, thereby reducing stress.

[0104] The telephone call support system can estimate the user's emotions and select background music during the call based on those emotions. For example, if the user is nervous, it can play relaxing music. If the user is relaxed, it can play music that promotes natural conversation. Furthermore, if the user is in a hurry, it can play music that encourages a quick response. This allows users to listen to background music that is optimal for their emotions while on a call, improving the quality of the call.

[0105] The telephone call support system can estimate the user's emotions and adjust the intensity of noise cancellation during the call based on those emotions. For example, if the user is nervous, strong noise cancellation can be applied to enhance their concentration. Conversely, if the user is relaxed, noise cancellation can be weakened to maintain natural speech. Furthermore, if the user is in a hurry, noise cancellation can be adjusted to facilitate a quick response. This allows users to utilize the optimal noise cancellation tailored to their emotions, thereby improving the quality of their calls.

[0106] The telephone call support system can translate conversations in real time, facilitating communication between users who speak different languages. For example, when a Japanese-speaking user and an English-speaking user are on a call, the conversation is translated in real time and conveyed to the other party. Similarly, translation can be performed when a French-speaking user and a Spanish-speaking user are on a call. Furthermore, even when users who speak different dialects are on a call, the dialect can be converted into standard Japanese. This facilitates smooth communication between users who speak different languages ​​or dialects.

[0107] A telephone call support system can automatically record what a user says during a call, allowing for later review. For example, it can record important statements during a business meeting for later review. It can also record important information during calls with family and friends. Furthermore, even if it's difficult to take notes during a call, the automatic recording function allows for later review. This ensures that users don't miss important information during calls and can review it later.

[0108] The telephone call support system can transcribe the user's speech in real time during a call, allowing for visual confirmation. For example, if a user with a hearing impairment is making a call, their speech will be transcribed and displayed on the screen. Similarly, even in noisy environments, the system can transcribe speech for confirmation. Furthermore, even if taking notes during a call is difficult, the real-time transcription function allows for visual confirmation of what is being said. This enables users to visually confirm what they are saying during a call, leading to more effective communication.

[0109] A telephone call support system can analyze a user's statements during a call and suggest appropriate responses. For example, it can suggest appropriate answers during a business meeting, supporting smooth progress. It can also suggest appropriate responses during customer support calls. Furthermore, it can suggest appropriate topics during calls with friends and family. This allows users to receive appropriate responses during calls, leading to more effective communication.

[0110] A telephone call support system can analyze user statements during a call and evaluate the call's progress in real time. For example, it can evaluate the progress of an agenda item during a business meeting and suggest the next steps. It can also evaluate the progress of problem-solving during a customer support call and suggest the next course of action. Furthermore, it can evaluate the progress of a conversation during a call with friends or family and suggest the next topic. This allows users to understand the progress of their calls in real time, enabling more effective communication.

[0111] The following briefly describes the processing flow for example form 2.

[0112] Step 1: The learning unit learns the user's voice. The user records their voice, and this recording data is input into the generating AI, which learns the characteristics of the user's voice. The generating AI learns features such as the tone, pitch, and intonation of the user's voice, and can generate speech that resembles the user's voice. Step 2: The reading unit reads the script based on the user's voice, which has been learned by the learning unit. When the user instructs the generating AI to read a script they have prepared in advance, the generating AI reads the script in the user's voice. The generating AI can produce natural-sounding speech while retaining the characteristics of the user's voice. Step 3: The adjustment unit adjusts the speaking speed of the audio read aloud by the reading unit. It adjusts the speaking speed based on the listener's reaction speed and speech clarity, and uses a generation AI to speak at an appropriate speed that is easy for the listener to understand.

[0113] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0114] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0115] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0116] Each of the multiple elements described above, including the learning unit, reading unit, adjustment unit, recording unit, input unit, and reference unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the learning unit is implemented by the processor 46 of the smart device 14, which records the user's voice and inputs the recorded data to the generating AI. The reading unit is implemented by the control unit 46A of the smart device 14, which reads the manuscript aloud in the user's voice. The adjustment unit is implemented by the specific processing unit 290 of the data processing unit 12, which adjusts the speaking speed. The recording unit is implemented by the microphone 38B of the smart device 14, which records the user's voice. The input unit inputs the recorded data to the generating AI via the communication I / F 44 of the smart device 14. The reference unit is implemented by the specific processing unit 290 of the data processing unit 12, which adjusts the speaking speed based on the other party's reaction speed and speech clarity. The correspondence between each unit and the device or control unit is not limited to the example described above, and various changes are possible.

[0117] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0118] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0119] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0120] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0121] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0122] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0123] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0124] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0125] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0126] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0127] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0128] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0129] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0130] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0131] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0132] Each of the multiple elements described above, including the learning unit, reading unit, adjustment unit, recording unit, input unit, and reference unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the learning unit is implemented by the processor 46 of the smart glasses 214, which records the user's voice and inputs the recorded data to the generating AI. The reading unit is implemented by the control unit 46A of the smart glasses 214, which reads the manuscript aloud in the user's voice. The adjustment unit is implemented by the specific processing unit 290 of the data processing unit 12, which adjusts the speaking speed. The recording unit is implemented by the microphone 238 of the smart glasses 214, which records the user's voice. The input unit inputs the recorded data to the generating AI via the communication I / F 44 of the smart glasses 214. The reference unit is implemented by the specific processing unit 290 of the data processing unit 12, which adjusts the speaking speed based on the other party's reaction speed and speech clarity. The correspondence between each unit and the device or control unit is not limited to the example described above, and various changes are possible.

[0133] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0134] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0135] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0136] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0137] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0138] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0139] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0140] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0141] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0142] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0143] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0144] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0145] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0146] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0147] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0148] Each of the multiple elements described above, including the learning unit, reading unit, adjustment unit, recording unit, input unit, and reference unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the learning unit is implemented by the processor 46 of the headset terminal 314, which records the user's voice and inputs the recorded data to the generating AI. The reading unit is implemented by the control unit 46A of the headset terminal 314, which reads the manuscript aloud in the user's voice. The adjustment unit is implemented by the specific processing unit 290 of the data processing unit 12, which adjusts the speaking speed. The recording unit is implemented by the microphone 238 of the headset terminal 314, which records the user's voice. The input unit inputs the recorded data to the generating AI via the communication I / F 44 of the headset terminal 314. The reference unit is implemented by the specific processing unit 290 of the data processing unit 12, which adjusts the speaking speed based on the other party's reaction speed and speech clarity. The correspondence between each unit and the device or control unit is not limited to the example described above, and various changes are possible.

[0149] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0150] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0151] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0152] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0153] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0154] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0155] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0156] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0157] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0158] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0159] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0160] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0161] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0162] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0163] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0164] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0165] Each of the multiple elements described above, including the learning unit, reading unit, adjustment unit, recording unit, input unit, and reference unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the learning unit is implemented by the processor 46 of the robot 414, which records the user's voice and inputs the recorded data to the generating AI. The reading unit is implemented by the control unit 46A of the robot 414, which reads the manuscript aloud in the user's voice. The adjustment unit is implemented by the specific processing unit 290 of the data processing unit 12, which adjusts the speaking speed. The recording unit is implemented by the microphone 238 of the robot 414, which records the user's voice. The input unit inputs the recorded data to the generating AI via the communication I / F 44 of the robot 414. The reference unit is implemented by the specific processing unit 290 of the data processing unit 12, which adjusts the speaking speed based on the other party's reaction speed and speech clarity. The correspondence between each unit and the device or control unit is not limited to the example described above, and various changes are possible.

[0166] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0167] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0168] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0169] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0170] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0171] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0172] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0173] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0174] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0175] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0176] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0177] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0178] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0179] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0180] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0181] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0182] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0183] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0184] (Note 1) A learning unit that learns from user feedback, A reading unit reads aloud a manuscript based on the user's voice learned by the aforementioned learning unit, The system includes an adjustment unit that adjusts the speaking speed of the audio read aloud by the aforementioned reading unit. A system characterized by the following features. (Note 2) It is equipped with a recording unit that records the user's voice. The system described in Appendix 1, characterized by the features described herein. (Note 3) The system includes an input unit that inputs the data recorded by the aforementioned recording unit into the generation AI. The system described in Appendix 2, characterized by the features described herein. (Note 4) The adjustment unit is, It features a reference unit that adjusts the speaking speed based on the other party's reaction speed and voice clarity. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned learning unit, The system estimates the user's emotions and selects training data based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned learning unit, During training, the system analyzes the user's voice tone and intonation in detail to generate more natural-sounding speech. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned learning unit, During learning, the system tracks changes in the user's voice in real time and reflects the latest voice characteristics. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned learning unit, It estimates the user's emotions and adjusts the learning frequency based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned learning unit, During training, background noise from the user's voice is removed to generate clear audio data. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned learning unit, During the learning process, the system learns different languages ​​and dialects from the user's voice to enable multilingual support. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned reading section is, It estimates the user's emotions and adjusts the tone and emotional expression of the reading based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned reading section is, When reading aloud, add appropriate intonation and emphasis according to the content of the script. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned reading section is, During narration, the system allows users to select different voice styles while preserving the characteristics of their own voice. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned reading section is, It estimates the user's emotions and adjusts the reading speed based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned reading section is, During reading aloud, appropriate pauses are inserted based on the content of the script to improve intelligibility. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned reading section is, During narration, echo and reverb effects are added to the user's voice to improve the sound quality. The system described in Appendix 1, characterized by the features described herein. (Note 17) The adjustment unit is, It estimates the user's emotions and adjusts the speaking speed based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 18) The adjustment unit is, During adjustments, the system monitors the other party's reaction speed in real time to maintain the optimal speaking speed. The system described in Appendix 1, characterized by the features described herein. (Note 19) The adjustment unit is, During adjustment, the system automatically evaluates the clarity of the audio and readjusts the speed as needed. The system described in Appendix 1, characterized by the features described herein. (Note 20) The adjustment unit is, It estimates the user's emotions and adjusts the volume based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 21) The adjustment unit is, During adjustment, the speaking speed is customized based on the other person's hearing characteristics. The system described in Appendix 1, characterized by the features described herein. (Note 22) The adjustment unit is, During adjustment, the system detects the level of background noise and applies noise cancellation to improve voice clarity. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned recording unit is It estimates the user's emotions and adjusts the recording start time based on the estimated emotions. The system described in Appendix 2, characterized by the features described herein. (Note 24) The aforementioned recording unit is During recording, the system automatically evaluates the quality of the user's voice and selects the optimal recording settings. The system described in Appendix 2, characterized by the features described herein. (Note 25) The aforementioned recording unit is It estimates the user's emotions and adjusts the length of the recording based on those emotions. The system described in Appendix 2, characterized by the features described herein. (Note 26) The aforementioned recording unit is During recording, multiple microphones are used to record spatial audio, generating more natural-sounding audio data. The system described in Appendix 2, characterized by the features described herein. (Note 27) The aforementioned input unit is It estimates the user's emotions and prioritizes input data based on the estimated user emotions. The system described in Appendix 3, characterized by the features described herein. (Note 28) The aforementioned input unit is During input, the system analyzes the characteristics of the user's voice in real time and selects the optimal input method. The system described in Appendix 3, characterized by the features described herein. (Note 29) The aforementioned input unit is It estimates the user's emotions and filters the input data based on the estimated user emotions. The system described in Appendix 3, characterized by the features described herein. (Note 30) The aforementioned input unit is During input, background noise from the user's voice is removed to generate clear audio data. The system described in Appendix 3, characterized by the features described herein. (Note 31) The aforementioned reference section is It estimates the user's emotions and adjusts the criteria based on the estimated user emotions. The system described in Appendix 4, characterized by the features described herein. (Note 32) The aforementioned reference section is When setting standards, the system monitors the other party's reaction speed and voice clarity in real time to maintain optimal standards. The system described in Appendix 4, characterized by the features described herein. (Note 33) The aforementioned reference section is It estimates the user's emotions and determines the priority of criteria based on the estimated user emotions. The system described in Appendix 4, characterized by the features described herein. (Note 34) The aforementioned reference section is When setting the criteria, customize the criteria based on the other person's hearing characteristics. The system described in Appendix 4, characterized by the features described herein. [Explanation of Symbols]

[0185] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A learning unit that learns from user feedback, A reading unit reads aloud a manuscript based on the user's voice learned by the aforementioned learning unit, The system includes an adjustment unit that adjusts the speaking speed of the audio read aloud by the aforementioned reading unit. A system characterized by the following features.

2. It is equipped with a recording unit that records the user's voice. The system according to feature 1.

3. The system includes an input unit that inputs the data recorded by the recording unit into the generation AI. The system according to feature 2.

4. The adjustment unit is, It features a reference unit that adjusts the speaking speed based on the other party's reaction speed and voice clarity. The system according to feature 1.

5. The aforementioned learning unit, The system estimates the user's emotions and selects training data based on those estimated emotions. The system according to feature 1.

6. The aforementioned learning unit, During training, the system analyzes the user's voice tone and intonation in detail to generate more natural-sounding speech. The system according to feature 1.

7. The aforementioned learning unit, During learning, the system tracks changes in the user's voice in real time and reflects the latest voice characteristics. The system according to feature 1.

8. The aforementioned learning unit, It estimates the user's emotions and adjusts the learning frequency based on the estimated user emotions. The system according to feature 1.

9. The aforementioned learning unit, During training, background noise from the user's voice is removed to generate clear audio data. The system according to feature 1.

10. The aforementioned learning unit, During the learning process, the system learns different languages ​​and dialects from the user's voice to enable multilingual support. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A