system
A system using lip-reading and eye-tracking technologies with generative AI generates natural-sounding speech and emotional intonation, addressing the challenge of voiceless communication, enhancing user experience and quality of life.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
People who cannot speak face challenges in communicating naturally due to their inability to produce voice.
A system utilizing lip-reading AI and eye-tracking technology for text input, combined with generative AI to generate natural-sounding speech and emotion generation to add intonation and emphasis, played back through devices like smartphones and PCs, enabling voice restoration.
Enables individuals who cannot speak to communicate naturally, improving their quality of life by allowing them to converse in their own voice, enhancing emotional expression and user convenience.
Smart Images

Figure 2026072327000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the prior art, there was a problem that it was difficult for people who cannot speak to communicate in natural voices.
[0005] The system according to the embodiment aims to enable people who cannot speak to communicate in natural voices.
Means for Solving the Problems
[0006] The system according to this embodiment comprises an input unit, a generation unit, an emotion generation unit, and a playback unit. The input unit receives text input using lip-reading AI or eye-tracking technology. The generation unit analyzes the text input by the input unit and generates natural speech using a generation AI that has learned the user's voice. The emotion generation unit adds intonation and emphasis to the speech generated by the generation unit. The playback unit plays back the speech generated by the emotion generation unit. [Effects of the Invention]
[0007] The system according to this embodiment allows people who are unable to speak to communicate using natural speech. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The system according to an embodiment of the present invention is a system that improves the quality of life (QOL) of people who are unable to speak due to illness. This system can simulate voice restoration with natural-sounding speech using a generative AI that has been trained on the user's voice. First, text input is performed using a device such as a smartphone, PC, AR, or MR glasses, either through lip-reading AI based on camera image analysis or eye-tracking via a keyboard. For example, the user's mouth movements are analyzed using the smartphone's camera, and the lip-reading AI generates text. Alternatively, the user selects a keyboard displayed on AR glasses with their gaze, and text is input using eye-tracking technology. Next, the generative AI analyzes the input text and generates natural-sounding speech using the generative AI trained on the user's voice. This generated speech is played back on a device such as a smartphone or PC. For example, if the user inputs "hello," the generative AI analyzes the text and plays back "hello" in the user's voice. Furthermore, the generated speech uses emotion generation engine technology to recognize emotions from facial expressions and add intonation and emphasis to the speech during conversation. For example, if the user inputs "thank you" with a smile, the emotion generation engine recognizes the facial expression and generates a voice with intonation that conveys gratitude. Furthermore, NFTs can be attached to the voice models of the generating AI to prove ownership. This guarantees that the generated voice belongs to the user. In addition, this system can be installed as an app on devices such as smartphones, tablets, and PCs, and can connect with AI speakers via Bluetooth® or other communication methods. This system can alleviate the sadness of losing one's voice and provide a more familiar environment for conversation partners by allowing them to converse in their own voice rather than a mechanical voice. For example, it can enable people who have lost their voice, such as those with ALS or laryngeal cancer, to communicate in their natural voice again. In this way, the system can improve the quality of life for people who are unable to speak due to illness.
[0029] The system according to this embodiment comprises an input unit, a generation unit, an emotion generation unit, and a playback unit. The input unit inputs text using lip-reading AI or eye-tracking technology. For example, the input unit analyzes the user's mouth movements using a smartphone camera, and the lip-reading AI generates the text. Alternatively, the input unit can also input text using eye-tracking technology after the user selects a character face displayed on AR glasses with their gaze. The generation unit analyzes the text input by the input unit and generates natural speech using a generation AI that has learned the user's voice. For example, the generation unit analyzes the input text using a generation AI and generates natural speech using the user's voice. Alternatively, the generation unit can also generate speech based on the content of the text using the generation AI. The emotion generation unit adds intonation and volume to the speech generated by the generation unit. For example, the emotion generation unit uses emotion generation engine technology to recognize emotions from facial expressions and adds intonation and volume to the speech during conversation. Alternatively, the emotion generation unit can use facial expression recognition technology to reflect emotions in the generated speech. The playback unit plays back the speech generated by the emotion generation unit. The playback unit plays audio on devices such as smartphones and PCs. The playback unit can also connect with AI speakers via Bluetooth or other communication methods to play audio. As a result, the system according to this embodiment can reproduce the user's voice in a natural way and play emotionally charged audio, thereby improving the quality of life (QOL) of people who are unable to speak.
[0030] The input unit uses lip-reading AI or eye-tracking technology to input text. Specifically, it analyzes the user's mouth movements using the smartphone's camera, and the lip-reading AI generates text. This lip-reading AI is trained using deep learning technology and can analyze various mouth movements and facial expressions with high accuracy. For example, when the user moves their mouth and pronounces "hello," the camera captures the movement, and the AI analyzes the movement to generate the corresponding text. The input unit can also input text by having the user select characters from a character face displayed on AR glasses using their gaze, and then using eye-tracking technology. The AR glasses display characters and symbols, and the user selects characters by moving their gaze. This eye-tracking technology tracks the user's gaze movements with high accuracy and can recognize selected characters in real time. This allows the user to input text using only their gaze, without using their hands. Furthermore, by combining these technologies, the input unit can provide flexible input methods tailored to the user's needs. For example, using lip-reading AI and eye-tracking technology together enables more accurate and faster text input. This allows the input unit to provide users with difficulty speaking with an efficient and intuitive means of text input, thereby facilitating smoother communication.
[0031] The generation unit analyzes the text input by the input unit and generates natural-sounding speech using a generation AI that has learned from the user's voice. Specifically, the generation AI analyzes the input text and generates natural-sounding speech using the user's voice. This generation AI learns the characteristics of the user's voice and can reproduce the tone, pitch, and rhythm of the voice. For example, if the user inputs "Good morning," the generation AI analyzes the text and generates natural-sounding speech using the user's voice. The generation unit can also generate speech based on the content of the text using the generation AI. For example, if the text is in the form of a question, the generation AI will generate speech with an intonation appropriate to that question. Furthermore, the generation unit can achieve a wider range of speech expressions by utilizing a database of user voices. For example, if the user wants to express different emotions, the generation AI can adjust the tone and rhythm of the voice according to those emotions. As a result, the generation unit can faithfully reproduce the user's voice and provide natural and diverse speech expressions. Moreover, because the generation unit can generate speech in real time, it can instantly generate speech in response to text input by the user, supporting smooth communication. This allows the voice generator to provide natural and diverse voice expressions to users who have difficulty speaking, thereby improving the quality of communication.
[0032] The emotion generation unit adds intonation and volume to the voice generated by the generation unit. Specifically, it uses emotion generation engine technology to recognize emotions from facial expressions and add intonation and volume to the voice during conversation. This emotion generation engine is trained using deep learning technology and can analyze the user's facial expressions and voice characteristics with high accuracy. For example, if the user is speaking with a smile, the emotion generation engine recognizes that expression and adds a bright and energetic tone to the generated voice. The emotion generation unit can also use facial recognition technology to reflect emotions in the generated voice. Facial recognition technology analyzes the user's facial expressions in real time through a camera and adjusts the intonation and volume of the voice based on the results. For example, if the user has a surprised expression, the emotion generation unit can recognize that expression and add a surprised tone to the generated voice. Furthermore, the emotion generation unit can achieve more accurate emotional expression by utilizing the user's past emotional data. For example, it can learn what kind of emotions the user has expressed in the past and adjust the intonation and volume of the voice based on those patterns. As a result, the emotion generation unit can generate natural voice that faithfully reflects the user's emotions, enabling richer communication.
[0033] The playback unit plays the audio generated by the emotion generation unit. Specifically, it plays the audio on devices such as smartphones and PCs. The playback unit employs the latest audio technology to reproduce the generated audio in high quality. For example, it can play clear and natural audio through smartphone speakers or earphones. The playback unit can also connect with AI speakers via Bluetooth or other communication methods to play audio. This allows users to play generated audio in various environments, such as at home or in the office. Furthermore, the playback unit has a function to adjust the playback speed and volume of the audio, allowing playback according to the user's preferences. For example, if a user wants to play the audio slowly, adjusting the playback speed provides easier listening. In addition, the playback unit can play audio in conjunction with multiple devices, so users are not dependent on a single device and can play audio on various devices. This allows the playback unit to provide a flexible audio playback environment that meets the user's needs and to play natural and emotionally rich audio for users who have difficulty speaking. Furthermore, the playback unit also has a function to record the audio playback history, allowing users to easily play audio they have played in the past. This allows the playback unit to improve user convenience and provide a more comfortable audio playback environment for users who have difficulty speaking.
[0034] The input unit can utilize devices such as smartphones, PCs, AR, and MR glasses. For example, the input unit can analyze the user's mouth movements using a smartphone camera, and a lip-reading AI can generate text. Alternatively, the input unit can allow the user to select a keyboard displayed on AR glasses using their gaze, and then input text using eye-tracking technology. Furthermore, the input unit can also be used with a PC or MR glasses, allowing the user to select a keyboard using their gaze and input text. This allows for improved user convenience through the use of a variety of devices.
[0035] The generation unit can utilize a generative AI trained on user voices. For example, the generation unit can analyze input text using the generative AI and generate natural-sounding speech using the user's voice. Furthermore, the generation unit can use the generative AI to generate speech based on the content of the text. In addition, the generation unit can use a large amount of audio data to train the generative AI on user voices. This allows for a natural reproduction of the user's voice.
[0036] The emotion generation unit can utilize technology to recognize emotions from facial expressions. For example, it can use emotion generation engine technology to recognize emotions from facial expressions and add intonation and volume to speech during conversation. The emotion generation unit can also use facial recognition technology to reflect emotions in the generated speech. Furthermore, the emotion generation unit can adjust the pitch and volume of speech to give it emotion. This makes it possible to generate speech that reflects the user's emotions.
[0037] The playback unit can play audio from devices such as smartphones and PCs. For example, it can play audio generated on devices such as smartphones and PCs. Furthermore, the playback unit can connect with AI speakers via Bluetooth or other communication methods to play audio. In addition, the playback unit has a function to adjust the sound quality when playing audio. This allows for audio playback on a variety of devices.
[0038] The voice model of the generating AI can be assigned an NFT to prove ownership. For example, blockchain technology can be used to assign an NFT to the voice model of the generating AI and prove ownership. In addition, specific identification information can be assigned to the voice model of the generating AI for ownership proof. Furthermore, a digital signature can be assigned to the voice data of the generating AI for ownership proof. This ensures that the generated voice belongs to the user.
[0039] The system can interact with AI speakers via communication methods such as Bluetooth. For example, the system can connect to an AI speaker via Bluetooth and play generated audio. The system can also interact with AI speakers using Wi-Fi or other wireless communication technologies. Furthermore, the system can provide a dedicated app for interacting with AI speakers. This expands the range of audio playback through interaction with AI speakers.
[0040] The input unit can analyze the user's past input history and select the optimal input method. For example, it can automatically display words and phrases that the user has frequently used in the past as suggestions. It can also prioritize suggesting input methods the user has used in the past (such as lip-reading or eye-tracking). Furthermore, it can predict and suggest words and phrases that the user will use at specific times based on their past input history. This allows the system to provide the optimal input method based on past input history.
[0041] The input unit can customize the input method based on the user's current health condition and environment. For example, if the user is tired, the input unit can provide a simpler input method. It can also prioritize voice input if the user is in a quiet environment. Furthermore, if the user is on the move, the input unit can provide an eye-tracking-based input method. This allows for input methods tailored to the user's health condition and environment.
[0042] The input unit can prioritize and provide the most relevant input method, taking into account the user's geographical location. For example, if the user is at home, the input unit can provide an input method suitable for a quiet environment. Furthermore, if the user is out and about, the input unit can provide an input method using eye tracking. Additionally, if the user is in a public place, the input unit can provide an input method using lip-reading. This allows the system to provide the optimal input method based on the user's geographical location.
[0043] The input unit can analyze a user's social media activity and provide relevant input methods. For example, it can analyze the content of social media posts frequently used by the user and suggest relevant words and phrases. It can also analyze the language and style used by the user on social media and provide the optimal input method. Furthermore, it can analyze the time of day when the user is active on social media and provide an input method suitable for that time. This allows for the provision of optimal input methods based on social media activity.
[0044] The generation unit can select the optimal voice generation method by referring to the user's past voice data during generation. For example, the generation unit can select the optimal voice generation method based on the voice data the user has used in the past. Furthermore, the generation unit can extract specific tones and pitches from the user's past voice data and incorporate them into voice generation. In addition, the generation unit can analyze the user's past voice data and select the most natural voice generation method. This allows the system to provide the optimal voice generation method based on past voice data.
[0045] The generation unit can customize the parameters of voice generation based on the user's current situation during generation. For example, if the user is in a quiet environment, the generation unit will generate clear voice. If the user is in a noisy environment, the generation unit can also generate voice with noise cancellation applied. Furthermore, if the user is moving, the generation unit can generate stable voice. This allows for the provision of an optimal voice generation method tailored to the current situation.
[0046] The generation unit can prioritize audio generation based on the user's submission timing. For example, if the user is in a hurry, the generation unit will prioritize generating the most important audio. Conversely, if the user is relaxed, the generation unit can also generate detailed audio. Furthermore, if the user needs audio at a specific time, the generation unit can generate audio to match that time. This allows for the provision of an optimal audio generation method based on the submission timing.
[0047] The generation unit can adjust the order of speech generation based on user relevance during the generation process. For example, it can prioritize generating phrases that the user frequently uses. It can also prioritize generating speech that the user uses in specific situations. Furthermore, it can prioritize generating the most relevant speech based on the user's past usage history. This allows for the provision of an optimal speech generation method based on relevance.
[0048] The emotion generation unit can select the optimal emotion generation method by referring to the user's past emotional data during emotion generation. For example, the emotion generation unit can select the optimal emotion generation method based on emotional data that the user has expressed in the past. Furthermore, the emotion generation unit can extract specific emotions from the user's past emotional data and reflect them in emotion generation. In addition, the emotion generation unit can analyze the user's past emotional data and select the most natural emotion generation method. This allows the system to provide the optimal emotion generation method based on past emotional data.
[0049] The emotion generation unit can customize the parameters of emotion generation based on the user's current facial expression. For example, if the user is smiling, the emotion generation unit will generate a cheerful emotion. It can also generate a calm emotion if the user has a serious expression. Furthermore, if the user is sad, the emotion generation unit can generate a comforting emotion. This allows for the provision of an optimal emotion generation method based on the user's current facial expression.
[0050] The emotion generation unit can select the optimal emotion generation method by considering the user's geographical location information during emotion generation. For example, if the user is at home, the emotion generation unit can generate a relaxed emotion. It can also generate an energetic emotion if the user is out and about. Furthermore, it can generate a subdued emotion if the user is in a public place. This allows for the provision of an optimal emotion generation method based on geographical location information.
[0051] The emotion generation unit can analyze the user's social media activity and propose methods for generating emotions. For example, it can analyze the content of social media posts that the user frequently uses and generate relevant emotions. It can also analyze the language and style used by the user on social media and provide the optimal emotion generation method. Furthermore, it can analyze the time of day when the user is active on social media and generate emotions appropriate for that time of day. This allows it to provide the optimal emotion generation method based on social media activity.
[0052] The playback unit can select the optimal playback method by referring to the user's past playback history during playback. For example, the playback unit can select the optimal playback method based on audio data that the user has played in the past. Furthermore, the playback unit can extract specific tones and pitches from the user's past playback history and reflect them in the playback. In addition, the playback unit can analyze the user's past playback history and select the most natural playback method. This allows the system to provide the optimal playback method based on past playback history.
[0053] The playback unit can customize the playback method based on the user's current environment during playback. For example, if the user is in a quiet environment, the playback unit will play back with clear audio. If the user is in a noisy environment, the playback unit can also play back with noise-canceling audio. Furthermore, if the user is on the move, the playback unit can play back with stable audio. This allows the system to provide the optimal playback method according to the current environment.
[0054] The playback unit can select the optimal playback method during playback, taking into account the user's geographical location. For example, if the user is at home, the playback unit will play in a relaxed voice. If the user is out, the playback unit can play in an energetic voice. Furthermore, if the user is in a public place, the playback unit can play in a quiet voice. This allows the system to provide an optimal playback method based on geographical location information.
[0055] The playback unit can analyze the user's social media activity during playback and suggest playback methods. For example, the playback unit can analyze the content of social media posts frequently used by the user and play relevant audio. It can also analyze the language and style used by the user on social media and provide the optimal playback method. Furthermore, the playback unit can analyze the user's social media activity times and play audio appropriate for those times. This allows for the provision of an optimal playback method based on social media activity.
[0056] The ownership verification unit of the voice model can select the optimal ownership verification method by referring to the user's past ownership data during ownership verification. For example, the ownership verification unit of the voice model can select the optimal ownership verification method based on data that the user has previously owned. Furthermore, the ownership verification unit of the voice model can extract specific verification methods from the user's past ownership data and reflect them in the ownership verification. In addition, the ownership verification unit of the voice model can analyze the user's past ownership data and select the most natural ownership verification method. This allows for the provision of the optimal ownership verification method based on past ownership data.
[0057] The voice model's ownership verification unit can select the optimal ownership verification method by considering the user's geographical location information during ownership verification. For example, if the user is at home, the voice model's ownership verification unit can provide a relaxed ownership verification method. It can also provide a quick ownership verification method if the user is out. Furthermore, if the user is in a public place, the voice model's ownership verification unit can provide a discreet ownership verification method. This allows for the provision of an optimal ownership verification method based on geographical location information.
[0058] The communication integration unit can select the optimal communication integration method by referring to the user's past communication history during communication integration. For example, the communication integration unit can select the optimal communication integration method based on the communication methods the user has used in the past. Furthermore, the communication integration unit can extract specific communication methods from the user's past communication history and reflect them in the communication integration. In addition, the communication integration unit can analyze the user's past communication history and select the most natural communication integration method. This allows the system to provide the optimal communication integration method based on past communication history.
[0059] The communication integration unit can select the optimal communication integration method when integrating with a user, taking into account the user's geographical location. For example, if the user is at home, the communication integration unit can provide a relaxed communication integration method. It can also provide a rapid communication integration method if the user is out and about. Furthermore, if the user is in a public place, the communication integration unit can provide a discreet communication integration method. This allows for the provision of the optimal communication integration method based on geographical location information.
[0060] The communication integration unit can analyze a user's social media activity and propose communication integration methods during communication integration. For example, the communication integration unit can analyze the content of social media posts frequently used by the user and provide relevant communication integration methods. It can also analyze the language and style used by the user on social media and provide the optimal communication integration method. Furthermore, the communication integration unit can analyze the user's social media activity times and provide communication integration methods suitable for those times. This allows for the provision of optimal communication integration methods based on social media activity.
[0061] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0062] The system can analyze a user's past input history and select the optimal input method. For example, it can automatically display words and phrases that the user has frequently used in the past as suggestions. It can also prioritize suggesting input methods that the user has used in the past (such as lip-reading or eye-tracking). Furthermore, it can predict and suggest words and phrases that the user will use at specific times based on their past input history. This allows the system to provide the optimal input method based on past input history.
[0063] The system can customize input methods based on the user's current health status and environment. For example, if the user is tired, it can provide a simpler input method. It can also prioritize voice input if the user is in a quiet environment. Furthermore, if the user is on the move, it can provide an input method using eye tracking. This allows the system to provide input methods tailored to the user's health status and environment.
[0064] The system can prioritize and provide the most relevant input methods by considering the user's geographical location. For example, if the user is at home, it can provide an input method suitable for a quiet environment. If the user is out, it can provide an input method using eye tracking. Furthermore, if the user is in a public place, it can provide an input method using lip-reading. This allows the system to provide the optimal input method based on the user's geographical location.
[0065] The system can analyze a user's social media activity and provide relevant input methods. For example, it can analyze the content of social media posts a user frequently uses and suggest relevant words and phrases. It can also analyze the language and style a user uses on social media and provide the optimal input method. Furthermore, it can analyze the time of day a user is active on social media and provide input methods appropriate for that time. This allows the system to provide the most suitable input method based on social media activity.
[0066] The system can select the optimal playback method by referring to the user's past playback history. For example, it can select the optimal playback method based on audio data the user has played in the past. It can also extract specific tones and pitches from the user's past playback history and reflect them in the playback. Furthermore, it can analyze the user's past playback history and select the most natural playback method. This allows the system to provide the optimal playback method based on past playback history.
[0067] The following briefly describes the processing flow for example form 1.
[0068] Step 1: The input unit uses lip-reading AI or eye-tracking technology to input text. For example, the user's mouth movements are analyzed using the smartphone camera, and the lip-reading AI generates the text. Alternatively, the user can select a character from a display on AR glasses using their gaze, and the text is entered using eye-tracking technology. Step 2: The generation unit analyzes the text input by the input unit and generates natural-sounding speech using a generation AI that has been trained on the user's voice. For example, the generation AI analyzes the input text and generates natural-sounding speech using the user's voice. The generation unit can also use the generation AI to generate speech based on the content of the text. Step 3: The emotion generation unit adds intonation and emphasis to the speech generated by the generation unit. For example, emotion generation engine technology can be used to recognize emotions from facial expressions and add intonation and emphasis to the speech during conversation. The emotion generation unit can also use facial recognition technology to reflect emotions in the generated speech. Step 4: The playback unit plays the audio generated by the emotion generation unit. For example, it plays the audio on a device such as a smartphone or PC. The playback unit can also connect with an AI speaker via Bluetooth or other communication methods to play the audio.
[0069] (Example of form 2) The system according to an embodiment of the present invention is a system that improves the quality of life (QOL) of people who are unable to speak due to illness. This system can simulate voice restoration with natural-sounding speech using a generative AI that has been trained on the user's voice. First, text input is performed using a device such as a smartphone, PC, AR, or MR glasses, either through lip-reading AI based on camera image analysis or eye-tracking via a keyboard. For example, the user's mouth movements are analyzed using the smartphone's camera, and the lip-reading AI generates text. Alternatively, the user selects a keyboard displayed on AR glasses with their gaze, and text is input using eye-tracking technology. Next, the generative AI analyzes the input text and generates natural-sounding speech using the generative AI trained on the user's voice. This generated speech is played back on a device such as a smartphone or PC. For example, if the user inputs "hello," the generative AI analyzes the text and plays back "hello" in the user's voice. Furthermore, the generated speech uses emotion generation engine technology to recognize emotions from facial expressions and add intonation and emphasis to the speech during conversation. For example, if the user inputs "thank you" with a smile, the emotion generation engine recognizes the facial expression and generates a voice with intonation that conveys gratitude. Furthermore, NFTs can be attached to the voice models of the generating AI to prove ownership. This ensures that the generated voice belongs to the user. In addition, this system can be installed as an app on devices such as smartphones, tablets, and PCs, and can connect with AI speakers via Bluetooth or other communication methods. This system can alleviate the sadness of losing one's voice and provide a more familiar environment for conversation partners by allowing them to converse in their own voice rather than a mechanical voice. For example, it can enable people who have lost their voice, such as those with ALS or laryngeal cancer, to communicate in their natural voice again. In this way, the system can improve the quality of life for people who are unable to speak due to illness.
[0070] The system according to this embodiment comprises an input unit, a generation unit, an emotion generation unit, and a playback unit. The input unit inputs text using lip-reading AI or eye-tracking technology. For example, the input unit analyzes the user's mouth movements using a smartphone camera, and the lip-reading AI generates the text. Alternatively, the input unit can also input text using eye-tracking technology after the user selects a character face displayed on AR glasses with their gaze. The generation unit analyzes the text input by the input unit and generates natural speech using a generation AI that has learned the user's voice. For example, the generation unit analyzes the input text using a generation AI and generates natural speech using the user's voice. Alternatively, the generation unit can also generate speech based on the content of the text using the generation AI. The emotion generation unit adds intonation and volume to the speech generated by the generation unit. For example, the emotion generation unit uses emotion generation engine technology to recognize emotions from facial expressions and adds intonation and volume to the speech during conversation. Alternatively, the emotion generation unit can use facial expression recognition technology to reflect emotions in the generated speech. The playback unit plays back the speech generated by the emotion generation unit. The playback unit plays audio on devices such as smartphones and PCs. The playback unit can also connect with AI speakers via Bluetooth or other communication methods to play audio. As a result, the system according to this embodiment can reproduce the user's voice in a natural way and play emotionally charged audio, thereby improving the quality of life (QOL) of people who are unable to speak.
[0071] The input unit uses lip-reading AI or eye-tracking technology to input text. Specifically, it analyzes the user's mouth movements using the smartphone's camera, and the lip-reading AI generates text. This lip-reading AI is trained using deep learning technology and can analyze various mouth movements and facial expressions with high accuracy. For example, when the user moves their mouth and pronounces "hello," the camera captures the movement, and the AI analyzes the movement to generate the corresponding text. The input unit can also input text by having the user select characters from a character face displayed on AR glasses using their gaze, and then using eye-tracking technology. The AR glasses display characters and symbols, and the user selects characters by moving their gaze. This eye-tracking technology tracks the user's gaze movements with high accuracy and can recognize selected characters in real time. This allows the user to input text using only their gaze, without using their hands. Furthermore, by combining these technologies, the input unit can provide flexible input methods tailored to the user's needs. For example, using lip-reading AI and eye-tracking technology together enables more accurate and faster text input. This allows the input unit to provide users with difficulty speaking with an efficient and intuitive means of text input, thereby facilitating smoother communication.
[0072] The generation unit analyzes the text input by the input unit and generates natural-sounding speech using a generation AI that has learned from the user's voice. Specifically, the generation AI analyzes the input text and generates natural-sounding speech using the user's voice. This generation AI learns the characteristics of the user's voice and can reproduce the tone, pitch, and rhythm of the voice. For example, if the user inputs "Good morning," the generation AI analyzes the text and generates natural-sounding speech using the user's voice. The generation unit can also generate speech based on the content of the text using the generation AI. For example, if the text is in the form of a question, the generation AI will generate speech with an intonation appropriate to that question. Furthermore, the generation unit can achieve a wider range of speech expressions by utilizing a database of user voices. For example, if the user wants to express different emotions, the generation AI can adjust the tone and rhythm of the voice according to those emotions. As a result, the generation unit can faithfully reproduce the user's voice and provide natural and diverse speech expressions. Moreover, because the generation unit can generate speech in real time, it can instantly generate speech in response to text input by the user, supporting smooth communication. This allows the voice generator to provide natural and diverse voice expressions to users who have difficulty speaking, thereby improving the quality of communication.
[0073] The emotion generation unit adds intonation and volume to the voice generated by the generation unit. Specifically, it uses emotion generation engine technology to recognize emotions from facial expressions and add intonation and volume to the voice during conversation. This emotion generation engine is trained using deep learning technology and can analyze the user's facial expressions and voice characteristics with high accuracy. For example, if the user is speaking with a smile, the emotion generation engine recognizes that expression and adds a bright and energetic tone to the generated voice. The emotion generation unit can also use facial recognition technology to reflect emotions in the generated voice. Facial recognition technology analyzes the user's facial expressions in real time through a camera and adjusts the intonation and volume of the voice based on the results. For example, if the user has a surprised expression, the emotion generation unit can recognize that expression and add a surprised tone to the generated voice. Furthermore, the emotion generation unit can achieve more accurate emotional expression by utilizing the user's past emotional data. For example, it can learn what kind of emotions the user has expressed in the past and adjust the intonation and volume of the voice based on those patterns. As a result, the emotion generation unit can generate natural voice that faithfully reflects the user's emotions, enabling richer communication.
[0074] The playback unit plays the audio generated by the emotion generation unit. Specifically, it plays the audio on devices such as smartphones and PCs. The playback unit employs the latest audio technology to reproduce the generated audio in high quality. For example, it can play clear and natural audio through smartphone speakers or earphones. The playback unit can also connect with AI speakers via Bluetooth or other communication methods to play audio. This allows users to play generated audio in various environments, such as at home or in the office. Furthermore, the playback unit has a function to adjust the playback speed and volume of the audio, allowing playback according to the user's preferences. For example, if a user wants to play the audio slowly, adjusting the playback speed provides easier listening. In addition, the playback unit can play audio in conjunction with multiple devices, so users are not dependent on a single device and can play audio on various devices. This allows the playback unit to provide a flexible audio playback environment that meets the user's needs and to play natural and emotionally rich audio for users who have difficulty speaking. Furthermore, the playback unit also has a function to record the audio playback history, allowing users to easily play audio they have played in the past. This allows the playback unit to improve user convenience and provide a more comfortable audio playback environment for users who have difficulty speaking.
[0075] The input unit can utilize devices such as smartphones, PCs, AR, and MR glasses. For example, the input unit can analyze the user's mouth movements using a smartphone camera, and a lip-reading AI can generate text. Alternatively, the input unit can allow the user to select a keyboard displayed on AR glasses using their gaze, and then input text using eye-tracking technology. Furthermore, the input unit can also be used with a PC or MR glasses, allowing the user to select a keyboard using their gaze and input text. This allows for improved user convenience through the use of a variety of devices.
[0076] The generation unit can utilize a generative AI trained on user voices. For example, the generation unit can analyze input text using the generative AI and generate natural-sounding speech using the user's voice. Furthermore, the generation unit can use the generative AI to generate speech based on the content of the text. In addition, the generation unit can use a large amount of audio data to train the generative AI on user voices. This allows for a natural reproduction of the user's voice.
[0077] The emotion generation unit can utilize technology to recognize emotions from facial expressions. For example, it can use emotion generation engine technology to recognize emotions from facial expressions and add intonation and volume to speech during conversation. The emotion generation unit can also use facial recognition technology to reflect emotions in the generated speech. Furthermore, the emotion generation unit can adjust the pitch and volume of speech to give it emotion. This makes it possible to generate speech that reflects the user's emotions.
[0078] The playback unit can play audio from devices such as smartphones and PCs. For example, it can play audio generated on devices such as smartphones and PCs. Furthermore, the playback unit can connect with AI speakers via Bluetooth or other communication methods to play audio. In addition, the playback unit has a function to adjust the sound quality when playing audio. This allows for audio playback on a variety of devices.
[0079] The voice model of the generating AI can be assigned an NFT to prove ownership. For example, blockchain technology can be used to assign an NFT to the voice model of the generating AI and prove ownership. In addition, specific identification information can be assigned to the voice model of the generating AI for ownership proof. Furthermore, a digital signature can be assigned to the voice data of the generating AI for ownership proof. This ensures that the generated voice belongs to the user.
[0080] The system can interact with AI speakers via communication methods such as Bluetooth. For example, the system can connect to an AI speaker via Bluetooth and play generated audio. The system can also interact with AI speakers using Wi-Fi or other wireless communication technologies. Furthermore, the system can provide a dedicated app for interacting with AI speakers. This expands the range of audio playback through interaction with AI speakers.
[0081] The input unit can estimate the user's emotions and adjust the input method based on the estimated emotions. For example, if the user is stressed, the input unit can provide a simple interface and minimize the input steps. If the user is relaxed, the input unit can also provide detailed input options and suggest a customizable input method. Furthermore, if the user is in a hurry, the input unit can prioritize voice input to allow for rapid text input. This allows for the provision of the optimal input method according to the user's emotions. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0082] The input unit can analyze the user's past input history and select the optimal input method. For example, it can automatically display words and phrases that the user has frequently used in the past as suggestions. It can also prioritize suggesting input methods the user has used in the past (such as lip-reading or eye-tracking). Furthermore, it can predict and suggest words and phrases that the user will use at specific times based on their past input history. This allows the system to provide the optimal input method based on past input history.
[0083] The input unit can customize the input method based on the user's current health condition and environment. For example, if the user is tired, the input unit can provide a simpler input method. It can also prioritize voice input if the user is in a quiet environment. Furthermore, if the user is on the move, the input unit can provide an eye-tracking-based input method. This allows for input methods tailored to the user's health condition and environment.
[0084] The input unit can estimate the user's emotions and determine the priority of inputs based on the estimated emotions. For example, if the user is nervous, the input unit will prioritize displaying important input items. It can also display detailed input items if the user is relaxed. Furthermore, if the user is in a hurry, the input unit can display only the most important input items. This provides input prioritization according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0085] The input unit can prioritize and provide the most relevant input method, taking into account the user's geographical location. For example, if the user is at home, the input unit can provide an input method suitable for a quiet environment. Furthermore, if the user is out and about, the input unit can provide an input method using eye tracking. Additionally, if the user is in a public place, the input unit can provide an input method using lip-reading. This allows the system to provide the optimal input method based on the user's geographical location.
[0086] The input unit can analyze a user's social media activity and provide relevant input methods. For example, it can analyze the content of social media posts frequently used by the user and suggest relevant words and phrases. It can also analyze the language and style used by the user on social media and provide the optimal input method. Furthermore, it can analyze the time of day when the user is active on social media and provide an input method suitable for that time. This allows for the provision of optimal input methods based on social media activity.
[0087] The generation unit can estimate the user's emotions and adjust the tone of the generated voice based on the estimated emotions. For example, if the user is relaxed, the generation unit will generate a calm tone of voice. It can also generate a lively tone of voice if the user is excited. Furthermore, if the user is sad, it can generate a calm tone of voice. This allows for the provision of a voice tone that corresponds to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generation AI. The generation AI may be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI.
[0088] The generation unit can select the optimal voice generation method by referring to the user's past voice data during generation. For example, the generation unit can select the optimal voice generation method based on the voice data the user has used in the past. Furthermore, the generation unit can extract specific tones and pitches from the user's past voice data and incorporate them into voice generation. In addition, the generation unit can analyze the user's past voice data and select the most natural voice generation method. This allows the system to provide the optimal voice generation method based on past voice data.
[0089] The generation unit can customize the parameters of voice generation based on the user's current situation during generation. For example, if the user is in a quiet environment, the generation unit will generate clear voice. If the user is in a noisy environment, the generation unit can also generate voice with noise cancellation applied. Furthermore, if the user is moving, the generation unit can generate stable voice. This allows for the provision of an optimal voice generation method tailored to the current situation.
[0090] The generation unit can estimate the user's emotions and adjust the length of the generated audio based on the estimated emotions. For example, if the user is in a hurry, the generation unit can generate a short, concise audio. If the user is relaxed, the generation unit can also generate a longer audio with detailed explanations. Furthermore, if the user is excited, the generation unit can generate audio with visually stimulating effects. This allows for audio lengths tailored to the user's emotions. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generation AI. Generation AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0091] The generation unit can prioritize audio generation based on the user's submission timing. For example, if the user is in a hurry, the generation unit will prioritize generating the most important audio. Conversely, if the user is relaxed, the generation unit can also generate detailed audio. Furthermore, if the user needs audio at a specific time, the generation unit can generate audio to match that time. This allows for the provision of an optimal audio generation method based on the submission timing.
[0092] The generation unit can adjust the order of speech generation based on user relevance during the generation process. For example, it can prioritize generating phrases that the user frequently uses. It can also prioritize generating speech that the user uses in specific situations. Furthermore, it can prioritize generating the most relevant speech based on the user's past usage history. This allows for the provision of an optimal speech generation method based on relevance.
[0093] The emotion generation unit can estimate the user's emotions and adjust the emotion generation method based on the estimated user emotions. For example, if the user is relaxed, the emotion generation unit can generate calm emotions. It can also generate lively emotions if the user is excited. Furthermore, if the user is sad, the emotion generation unit can generate calm emotions. This allows for the provision of an optimal emotion generation method tailored to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generation AI. The generation AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0094] The emotion generation unit can select the optimal emotion generation method by referring to the user's past emotional data during emotion generation. For example, the emotion generation unit can select the optimal emotion generation method based on emotional data that the user has expressed in the past. Furthermore, the emotion generation unit can extract specific emotions from the user's past emotional data and reflect them in emotion generation. In addition, the emotion generation unit can analyze the user's past emotional data and select the most natural emotion generation method. This allows the system to provide the optimal emotion generation method based on past emotional data.
[0095] The emotion generation unit can customize the parameters of emotion generation based on the user's current facial expression. For example, if the user is smiling, the emotion generation unit will generate a cheerful emotion. It can also generate a calm emotion if the user has a serious expression. Furthermore, if the user is sad, the emotion generation unit can generate a comforting emotion. This allows for the provision of an optimal emotion generation method based on the user's current facial expression.
[0096] The emotion generation unit can estimate the user's emotions and determine the priority of emotion generation based on the estimated user emotions. For example, if the user is nervous, the emotion generation unit will prioritize generating important emotions. It can also generate detailed emotions if the user is relaxed. Furthermore, if the user is in a hurry, the emotion generation unit can generate only the most important emotions. This allows for prioritizing emotion generation according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generation AI. The generation AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0097] The emotion generation unit can select the optimal emotion generation method by considering the user's geographical location information during emotion generation. For example, if the user is at home, the emotion generation unit can generate a relaxed emotion. It can also generate an energetic emotion if the user is out and about. Furthermore, it can generate a subdued emotion if the user is in a public place. This allows for the provision of an optimal emotion generation method based on geographical location information.
[0098] The emotion generation unit can analyze the user's social media activity and propose methods for generating emotions. For example, it can analyze the content of social media posts that the user frequently uses and generate relevant emotions. It can also analyze the language and style used by the user on social media and provide the optimal emotion generation method. Furthermore, it can analyze the time of day when the user is active on social media and generate emotions appropriate for that time of day. This allows it to provide the optimal emotion generation method based on social media activity.
[0099] The playback unit can estimate the user's emotions and adjust the playback method based on the estimated emotions. For example, if the user is relaxed, the playback unit will play in a calm voice. If the user is excited, the playback unit can play in a lively voice. Furthermore, if the user is sad, the playback unit can play in a calm voice. This allows the system to provide the optimal playback method according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0100] The playback unit can select the optimal playback method by referring to the user's past playback history during playback. For example, the playback unit can select the optimal playback method based on audio data that the user has played in the past. Furthermore, the playback unit can extract specific tones and pitches from the user's past playback history and reflect them in the playback. In addition, the playback unit can analyze the user's past playback history and select the most natural playback method. This allows the system to provide the optimal playback method based on past playback history.
[0101] The playback unit can customize the playback method based on the user's current environment during playback. For example, if the user is in a quiet environment, the playback unit will play back with clear audio. If the user is in a noisy environment, the playback unit can also play back with noise-canceling audio. Furthermore, if the user is on the move, the playback unit can play back with stable audio. This allows the system to provide the optimal playback method according to the current environment.
[0102] The playback unit can estimate the user's emotions and determine playback priorities based on those estimated emotions. For example, if the user is tense, the playback unit will prioritize playing important audio. Conversely, if the user is relaxed, the playback unit can also play detailed audio. Furthermore, if the user is in a hurry, the playback unit can play only the most important audio. This allows for playback prioritization tailored to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0103] The playback unit can select the optimal playback method during playback, taking into account the user's geographical location. For example, if the user is at home, the playback unit will play in a relaxed voice. If the user is out, the playback unit can play in an energetic voice. Furthermore, if the user is in a public place, the playback unit can play in a quiet voice. This allows the system to provide an optimal playback method based on geographical location information.
[0104] The playback unit can analyze the user's social media activity during playback and suggest playback methods. For example, the playback unit can analyze the content of social media posts frequently used by the user and play relevant audio. It can also analyze the language and style used by the user on social media and provide the optimal playback method. Furthermore, the playback unit can analyze the user's social media activity times and play audio appropriate for those times. This allows for the provision of an optimal playback method based on social media activity.
[0105] The ownership verification unit of the voice model can estimate the user's emotions and adjust the ownership verification method based on the estimated emotions. For example, if the user is relaxed, the ownership verification unit of the voice model can provide a simple ownership verification method. If the user is tense, it can also provide a more detailed ownership verification method. Furthermore, if the user is in a hurry, it can provide a rapid ownership verification method. This allows for the provision of the optimal ownership verification method according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0106] The ownership verification unit of the voice model can select the optimal ownership verification method by referring to the user's past ownership data during ownership verification. For example, the ownership verification unit of the voice model can select the optimal ownership verification method based on data that the user has previously owned. Furthermore, the ownership verification unit of the voice model can extract specific verification methods from the user's past ownership data and reflect them in the ownership verification. In addition, the ownership verification unit of the voice model can analyze the user's past ownership data and select the most natural ownership verification method. This allows for the provision of the optimal ownership verification method based on past ownership data.
[0107] The ownership verification unit of the voice model can estimate the user's emotions and determine the priority of ownership verification based on the estimated emotions. For example, if the user is nervous, the ownership verification unit of the voice model will prioritize important ownership verification. If the user is relaxed, the ownership verification unit of the voice model can also perform detailed ownership verification. Furthermore, if the user is in a hurry, the ownership verification unit of the voice model can perform only the most important ownership verification. This allows for prioritization of ownership verification according to the user's emotions. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0108] The voice model's ownership verification unit can select the optimal ownership verification method by considering the user's geographical location information during ownership verification. For example, if the user is at home, the voice model's ownership verification unit can provide a relaxed ownership verification method. It can also provide a quick ownership verification method if the user is out. Furthermore, if the user is in a public place, the voice model's ownership verification unit can provide a discreet ownership verification method. This allows for the provision of an optimal ownership verification method based on geographical location information.
[0109] The communication unit can estimate the user's emotions and adjust the communication method based on the estimated emotions. For example, if the user is relaxed, the communication unit can provide a calm communication method. If the user is excited, the communication unit can also provide a rapid communication method. Furthermore, if the user is sad, the communication unit can also provide a calm communication method. This allows for the provision of the optimal communication method according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0110] The communication integration unit can select the optimal communication integration method by referring to the user's past communication history during communication integration. For example, the communication integration unit can select the optimal communication integration method based on the communication methods the user has used in the past. Furthermore, the communication integration unit can extract specific communication methods from the user's past communication history and reflect them in the communication integration. In addition, the communication integration unit can analyze the user's past communication history and select the most natural communication integration method. This allows the system to provide the optimal communication integration method based on past communication history.
[0111] The communication coordination unit can estimate the user's emotions and determine the priority of communication coordination based on the estimated emotions. For example, if the user is stressed, the communication coordination unit will prioritize important communication coordination. If the user is relaxed, the communication coordination unit can also prioritize detailed communication coordination. Furthermore, if the user is in a hurry, the communication coordination unit can prioritize only the most important communication coordination. This allows for the provision of communication coordination priorities that correspond to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0112] The communication integration unit can select the optimal communication integration method when integrating with a user, taking into account the user's geographical location. For example, if the user is at home, the communication integration unit can provide a relaxed communication integration method. It can also provide a rapid communication integration method if the user is out and about. Furthermore, if the user is in a public place, the communication integration unit can provide a discreet communication integration method. This allows for the provision of the optimal communication integration method based on geographical location information.
[0113] The communication integration unit can analyze a user's social media activity and propose communication integration methods during communication integration. For example, the communication integration unit can analyze the content of social media posts frequently used by the user and provide relevant communication integration methods. It can also analyze the language and style used by the user on social media and provide the optimal communication integration method. Furthermore, the communication integration unit can analyze the user's social media activity times and provide communication integration methods suitable for those times. This allows for the provision of optimal communication integration methods based on social media activity.
[0114] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0115] The system can estimate the user's emotions and adjust the tone of voice based on those emotions. For example, if the user is relaxed, it can generate a calm tone of voice. If the user is excited, it can generate a lively tone of voice. Furthermore, if the user is sad, it can generate a calm tone of voice. This allows the system to provide a tone of voice that matches the user's emotions.
[0116] The system can analyze a user's past input history and select the optimal input method. For example, it can automatically display words and phrases that the user has frequently used in the past as suggestions. It can also prioritize suggesting input methods that the user has used in the past (such as lip-reading or eye-tracking). Furthermore, it can predict and suggest words and phrases that the user will use at specific times based on their past input history. This allows the system to provide the optimal input method based on past input history.
[0117] The system can estimate the user's emotions and adjust the input method based on those emotions. For example, if the user is stressed, it can provide a simple interface and minimize the input steps. If the user is relaxed, it can provide detailed input options and suggest a customizable input method. Furthermore, if the user is in a hurry, it can prioritize voice input to allow for quick text entry. This allows the system to provide the optimal input method according to the user's emotions.
[0118] The system can customize input methods based on the user's current health status and environment. For example, if the user is tired, it can provide a simpler input method. It can also prioritize voice input if the user is in a quiet environment. Furthermore, if the user is on the move, it can provide an input method using eye tracking. This allows the system to provide input methods tailored to the user's health status and environment.
[0119] The system can estimate the user's emotions and adjust the length of the generated audio based on those emotions. For example, if the user is in a hurry, it can generate a short, concise audio. If the user is relaxed, it can generate a longer audio with more detailed explanations. Furthermore, if the user is excited, it can generate audio with visually stimulating effects. This allows the system to provide audio lengths that match the user's emotions.
[0120] The system can prioritize and provide the most relevant input methods by considering the user's geographical location. For example, if the user is at home, it can provide an input method suitable for a quiet environment. If the user is out, it can provide an input method using eye tracking. Furthermore, if the user is in a public place, it can provide an input method using lip-reading. This allows the system to provide the optimal input method based on the user's geographical location.
[0121] The system can estimate the user's emotions and adjust its emotion generation method based on the estimated emotions. For example, if the user is relaxed, it can generate calm emotions. If the user is excited, it can generate lively emotions. Furthermore, if the user is sad, it can generate calm emotions. This allows the system to provide the optimal emotion generation method according to the user's emotions.
[0122] The system can analyze a user's social media activity and provide relevant input methods. For example, it can analyze the content of social media posts a user frequently uses and suggest relevant words and phrases. It can also analyze the language and style a user uses on social media and provide the optimal input method. Furthermore, it can analyze the time of day a user is active on social media and provide input methods appropriate for that time. This allows the system to provide the most suitable input method based on social media activity.
[0123] The system can estimate the user's emotions and adjust the playback method based on those estimates. For example, if the user is relaxed, it can play in a calm voice. If the user is excited, it can play in an energetic voice. Furthermore, if the user is sad, it can play in a soothing voice. This allows the system to provide the optimal playback method according to the user's emotions.
[0124] The system can select the optimal playback method by referring to the user's past playback history. For example, it can select the optimal playback method based on audio data the user has played in the past. It can also extract specific tones and pitches from the user's past playback history and reflect them in the playback. Furthermore, it can analyze the user's past playback history and select the most natural playback method. This allows the system to provide the optimal playback method based on past playback history.
[0125] The following briefly describes the processing flow for example form 2.
[0126] Step 1: The input unit uses lip-reading AI or eye-tracking technology to input text. For example, the user's mouth movements are analyzed using the smartphone camera, and the lip-reading AI generates the text. Alternatively, the user can select a character from a display on AR glasses using their gaze, and the text is entered using eye-tracking technology. Step 2: The generation unit analyzes the text input by the input unit and generates natural-sounding speech using a generation AI that has been trained on the user's voice. For example, the generation AI analyzes the input text and generates natural-sounding speech using the user's voice. The generation unit can also use the generation AI to generate speech based on the content of the text. Step 3: The emotion generation unit adds intonation and emphasis to the speech generated by the generation unit. For example, emotion generation engine technology can be used to recognize emotions from facial expressions and add intonation and emphasis to the speech during conversation. The emotion generation unit can also use facial recognition technology to reflect emotions in the generated speech. Step 4: The playback unit plays the audio generated by the emotion generation unit. For example, it plays the audio on a device such as a smartphone or PC. The playback unit can also connect with an AI speaker via Bluetooth or other communication methods to play the audio.
[0127] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0128] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0129] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0130] Each of the multiple elements described above, including the input unit, generation unit, emotion generation unit, and playback unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the input unit uses the camera 42 of the smart device 14 to analyze the user's mouth movements, and a lip-reading AI generates text. Alternatively, the input unit can use the control unit 46A of the smart device 14 to select a character face displayed on AR glasses with the user's gaze and input text using eye-tracking technology. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12, for example, to analyze the input text and generate natural-sounding speech using the user's voice. The emotion generation unit is implemented in the specific processing unit 290 of the data processing unit 12, for example, to recognize emotions from facial expressions using emotion generation engine technology and add intonation and emphasis to the generated speech. The playback unit plays back the speech using the output device 40 of the smart device 14, for example. Alternatively, the playback unit can play back the speech by coordinating with the smart device 14 and an AI speaker via communication such as Bluetooth. The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.
[0131] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0132] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0133] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0134] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0135] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0136] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0137] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0138] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0139] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0140] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0141] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0142] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0143] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0144] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0145] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0146] Each of the multiple elements described above, including the input unit, generation unit, emotion generation unit, and playback unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the input unit uses the camera 42 of the smart glasses 214 to analyze the user's mouth movements, and a lip-reading AI generates text. Alternatively, the input unit can use the control unit 46A of the smart glasses 214 to select a character face displayed on the AR glasses with the user's gaze and input text using eye-tracking technology. The generation unit is implemented, for example, by the specific processing unit 290 of the data processing unit 12, which analyzes the input text and generates natural-sounding speech using the user's voice. The emotion generation unit is implemented, for example, by the specific processing unit 290 of the data processing unit 12, which uses emotion generation engine technology to recognize emotions from facial expressions and add intonation and intensity to the generated speech. The playback unit plays back speech using, for example, the speaker 240 of the smart glasses 214. Alternatively, the playback unit can play back speech by coordinating the smart glasses 214 and the AI speaker via communication such as Bluetooth. The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.
[0147] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0148] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0149] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0150] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0151] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0152] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0153] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0154] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0155] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0156] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0157] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0158] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0159] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0160] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0161] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0162] Each of the multiple elements described above, including the input unit, generation unit, emotion generation unit, and playback unit, is implemented in at least one of the following: the headset terminal 314 and the data processing unit 12. For example, the input unit uses the camera 42 of the headset terminal 314 to analyze the user's mouth movements, and a lip-reading AI generates text. Alternatively, the input unit can use the control unit 46A of the headset terminal 314 to select a character face displayed on AR glasses with the user's gaze and input text using eye-tracking technology. The generation unit is implemented, for example, by the specific processing unit 290 of the data processing unit 12, which analyzes the input text and generates natural-sounding speech using the user's voice. The emotion generation unit is implemented, for example, by the specific processing unit 290 of the data processing unit 12, which uses emotion generation engine technology to recognize emotions from facial expressions and add intonation and emphasis to the generated speech. The playback unit plays back speech using, for example, the speaker 240 of the headset terminal 314. Alternatively, the playback unit can play back speech by coordinating the headset terminal 314 and the AI speaker via communication such as Bluetooth. The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.
[0163] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0164] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0165] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0166] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0167] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0168] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0169] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0170] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0171] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0172] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0173] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0174] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0175] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0176] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0177] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0178] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0179] Each of the multiple elements described above, including the input unit, generation unit, emotion generation unit, and playback unit, is implemented in at least one of the following: the robot 414 and the data processing unit 12. For example, the input unit uses the camera 42 of the robot 414 to analyze the user's mouth movements, and a lip-reading AI generates text. Alternatively, the input unit can use the control unit 46A of the robot 414 to select a character face displayed on AR glasses with the user's gaze and input text using eye-tracking technology. The generation unit is implemented, for example, by the specific processing unit 290 of the data processing unit 12, which analyzes the input text and generates natural-sounding speech using the user's voice. The emotion generation unit is implemented, for example, by the specific processing unit 290 of the data processing unit 12, which uses emotion generation engine technology to recognize emotions from facial expressions and add intonation and emphasis to the generated speech. The playback unit plays back speech using, for example, the speaker 240 of the robot 414. Alternatively, the playback unit can play back speech through communication such as Bluetooth between the robot 414 and the AI speaker. The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.
[0180] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0181] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0182] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0183] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0184] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0185] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0186] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0187] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0188] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0189] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0190] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0191] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0192] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0193] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0194] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0195] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0196] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0197] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0198] (Note 1) An input unit that uses lip-reading AI or eye-tracking technology to input text, A generation unit analyzes the text input by the aforementioned input unit and generates natural speech using a generation AI that has learned the user's voice. An emotion generation unit that adds intonation and intensity to the sound generated by the generation unit, The system includes a playback unit that plays back the sound generated by the emotion generation unit. A system characterized by the following features. (Note 2) The aforementioned input unit is Using devices such as smartphones, PCs, AR, and MR glasses The system described in Appendix 1, characterized by the features described herein. (Note 3) The generating unit is Uses a generative AI trained on user feedback. The system described in Appendix 1, characterized by the features described herein. (Note 4) The emotion generation unit, This technology uses facial expressions to recognize emotions. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned regeneration unit is Play audio on devices such as smartphones and PCs. The system described in Appendix 1, characterized by the features described herein. (Note 6) The speech model of the generative AI includes: Assigning an NFT to prove ownership The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned system, Connects with AI speakers via Bluetooth and other communication methods. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned input unit is It estimates the user's emotions and adjusts the input method based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned input unit is Analyze the user's past input history and select the optimal input method. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned input unit is Customize input methods based on the user's current health status and environment. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned input unit is It estimates the user's emotions and determines the priority of inputs based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned input unit is Prioritize providing the most relevant input method, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned input unit is Analyze users' social media activity and provide relevant input methods. The system described in Appendix 1, characterized by the features described herein. (Note 14) The generating unit is It estimates the user's emotions and adjusts the tone of the generated voice based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 15) The generating unit is During generation, the system references the user's past voice data to select the optimal voice generation method. The system described in Appendix 1, characterized by the features described herein. (Note 16) The generating unit is During generation, the parameters for voice generation are customized based on the user's current situation. The system described in Appendix 1, characterized by the features described herein. (Note 17) The generating unit is It estimates the user's emotions and adjusts the length of the generated audio based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 18) The generating unit is During generation, the priority of voice generation is determined based on when the user submitted their work. The system described in Appendix 1, characterized by the features described herein. (Note 19) The generating unit is During generation, the order of speech generation is adjusted based on user relevance. The system described in Appendix 1, characterized by the features described herein. (Note 20) The emotion generation unit, It estimates the user's emotions and adjusts the emotion generation method based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 21) The emotion generation unit, When generating emotions, the system selects the optimal emotion generation method by referring to the user's past emotional data. The system described in Appendix 1, characterized by the features described herein. (Note 22) The emotion generation unit, When generating emotions, the parameters for emotion generation are customized based on the user's current facial expression. The system described in Appendix 1, characterized by the features described herein. (Note 23) The emotion generation unit, It estimates the user's emotions and determines the priority of emotion generation based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 24) The emotion generation unit, When generating emotions, the system selects the optimal emotion generation method by considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 25) The emotion generation unit, When generating emotions, we analyze users' social media activity and propose methods for generating those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned regeneration unit is It estimates the user's emotions and adjusts the playback method based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned regeneration unit is During playback, the system selects the optimal playback method by referring to the user's past playback history. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned regeneration unit is During playback, the playback method is customized based on the user's current environment. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned regeneration unit is It estimates the user's emotions and determines playback priorities based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned regeneration unit is During playback, the system selects the optimal playback method by considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned regeneration unit is During playback, the system analyzes the user's social media activity and suggests playback methods. The system described in Appendix 1, characterized by the features described herein. (Note 32) The ownership verification unit of the aforementioned voice model is: The system estimates the user's emotions and adjusts the ownership verification method based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 33) The ownership verification unit of the aforementioned voice model is: During ownership verification, the system selects the most suitable ownership verification method by referring to the user's past ownership data. The system described in Appendix 1, characterized by the features described herein. (Note 34) The ownership verification unit of the aforementioned voice model is: The system estimates the user's emotions and determines the priority of ownership verification based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 35) The ownership verification unit of the aforementioned voice model is: When verifying ownership, the system selects the most suitable ownership verification method, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 36) The aforementioned communication cooperation unit is It estimates the user's emotions and adjusts the communication method based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 37) The aforementioned communication cooperation unit is When establishing a communication connection, the system selects the optimal communication method by referring to the user's past communication history. The system described in Appendix 1, characterized by the features described herein. (Note 38) The aforementioned communication cooperation unit is It estimates the user's emotions and determines the priority of communication interactions based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 39) The aforementioned communication cooperation unit is When establishing a communication connection, the system selects the optimal communication method by considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 40) The aforementioned communication cooperation unit is When integrating communications, we analyze the user's social media activity and propose methods for communication integration. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0199] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. An input unit that uses lip-reading AI or eye-tracking technology to input text, A generation unit analyzes the text input by the aforementioned input unit and generates natural-sounding speech using a generation AI that has learned the user's voice. An emotion generation unit that adds intonation and intensity to the sound generated by the generation unit, The system includes a playback unit that plays back the sound generated by the emotion generation unit. A system characterized by the following features.
2. The aforementioned input unit is Using devices such as smartphones, PCs, AR, and MR glasses The system according to feature 1.
3. The generating unit is Uses generative AI trained on user feedback. The system according to feature 1.
4. The emotion generation unit, This technology uses facial expressions to recognize emotions. The system according to feature 1.
5. The aforementioned regeneration unit is Play audio on devices such as smartphones and PCs. The system according to feature 1.
6. The speech model for the generative AI includes: Assigning an NFT to prove ownership The system according to feature 1.
7. The aforementioned system, Connects with AI speakers via communication. The system according to feature 1.
8. The aforementioned input unit is It estimates the user's emotions and adjusts the input method based on the estimated emotions. The system according to feature 1.
9. The aforementioned input unit is Analyze the user's past input history and select the optimal input method. The system according to feature 1.
10. The aforementioned input unit is Customize input methods based on the user's current health status and environment. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A