system

The system addresses the challenge of individuals with speech disorders by learning from past data and family voices to generate personalized voices, enhancing communication quality and naturalness.

JP2026072476APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

People with speech disorders face challenges in speaking in a voice that closely resembles their own voice.

Method used

A system comprising a reception unit, a learning unit, and a voice generation unit that learns from past recording data and family voices to generate voices that reflect individuality, using augmented reality to enhance face-to-face communication.

Benefits of technology

Enables individuals with speech impairments to converse using a voice that closely resembles their own, improving the quality and naturalness of communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072476000001_ABST
    Figure 2026072476000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to enable people with speech impairments to converse using a voice that is close to their own. [Solution] The system according to the embodiment comprises a reception unit, a learning unit, a voice generation unit, and an output unit. The reception unit receives user input. The learning unit learns past recording data and family voices based on the information received by the reception unit. The voice generation unit generates voices that reflect individuality based on the data learned by the learning unit. The output unit outputs the voices generated by the voice generation unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] ,

[0006] , , , , , ,

[0005] , , ,

[0003] , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, there was a problem that it was difficult for people with speech disorders to talk in a voice close to their own voice.

[0005] The system according to the embodiment aims to enable people with speech disorders to talk in a voice close to their own voice.

Means for Solving the Problems

[0006] The system according to this embodiment comprises a reception unit, a learning unit, a voice generation unit, and an output unit. The reception unit receives user input. The learning unit learns past recording data and family voices based on the information received by the reception unit. The voice generation unit generates voices that reflect individuality based on the data learned by the learning unit. The output unit outputs the voices generated by the voice generation unit. [Effects of the Invention]

[0007] The system according to this embodiment can enable people with speech impairments to converse using a voice that is close to their own. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters linked by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The speech support application according to an embodiment of the present invention is a novel speech support application that enables people with speech disorders to converse in a voice that is closer to their own voice. This speech support application learns the user's past recording data and the voices of family members to achieve speech synthesis that reflects the user's individuality. Specifically, it uses a generation AI to learn from the user's and their family's past recording data and generates a voice that reflects their individuality. Furthermore, by adding augmented reality (AR) functionality, it enhances face-to-face communication and naturally reflects the other person's facial expressions and gestures. For example, the speech support application collects the user's past recording data, and the generation AI learns from that data. The generation AI extracts the characteristics of the user's voice and generates speech based on that. In addition, the speech support application uses AR functionality to recognize the facial expressions and gestures of the person it is talking to in real time and provides feedback to the user. As a result, the user can have a natural conversation in a voice that reflects their individuality, improving the quality of communication. Moreover, the speech support application can be used on smartphones and AR glasses, improving the user's quality of life and increasing opportunities for social participation. This allows speech support apps to enable natural conversations using voices that reflect the user's individuality, thereby improving the quality of communication.

[0029] The speech support application according to this embodiment comprises a reception unit, a learning unit, a speech generation unit, and an output unit. The reception unit receives user input. User input includes, but is not limited to, voice input, text input, and gesture input. For example, the reception unit receives voice input via a microphone. The reception unit can also receive text input via a keyboard or touchscreen. Furthermore, the reception unit can also receive gesture input via a camera or sensor. For example, the reception unit receives voice input via a microphone and converts it to text using speech recognition technology. Text input can be entered directly using a keyboard or touchscreen. Gesture input is recognized by detecting the user's movements using a camera or sensor. The learning unit learns past recording data and family voices based on the information received by the reception unit. For example, the learning unit collects the user's past recording data and learns that data using a generation AI. For example, the learning unit extracts the characteristics of the user's voice and generates speech based on them. The learning unit can also collect the voices of family members and learn that data using a generation AI. For example, the learning unit collects the user's past recorded data and inputs it into the generating AI. The generating AI extracts the characteristics of the user's voice and generates speech based on that. Family members' voices are also collected in a similar manner and input into the generating AI for training. The speech generation unit generates speech that reflects individuality based on the data learned by the learning unit. For example, the speech generation unit uses the generating AI to generate speech based on the characteristics of the user's voice. For example, the speech generation unit generates speech that reflects the tone, pitch, and accent of the user's voice using the generating AI. The speech generation unit can also generate speech based on the characteristics of family members' voices using the generating AI. For example, the speech generation unit generates speech that reflects the tone, pitch, and accent of the user's voice using the generating AI. Family members' voices are also generated using the generating AI in a similar manner. The output unit outputs the speech generated by the speech generation unit. For example, the output unit outputs speech using a speaker. For example, the output unit outputs speech using a speaker. The output unit can also output speech using headphones or earphones. Furthermore, the output unit can also output speech using AR glasses.For example, the output unit outputs sound using a speaker. It can also output sound using headphones or earphones. It can also output sound using AR glasses. As a result, the speech support application according to the embodiment can achieve natural dialogue with a voice that reflects the user's individuality, thereby improving the quality of communication.

[0030] The reception unit receives user input. User input includes, but is not limited to, voice input, text input, and gesture input. For example, the reception unit can receive voice input via a microphone. It can also receive text input via a keyboard or touchscreen. Furthermore, it can receive gesture input via a camera or sensor. For example, the reception unit can receive voice input via a microphone and convert it to text using speech recognition technology. Text input can be done directly using a keyboard or touchscreen. Gesture input detects user movements using a camera or sensor and recognizes them as input. The reception unit comprehensively manages these diverse input methods and provides an interface to accurately understand the user's intent. For example, in the case of voice input, a highly sensitive microphone with noise cancellation is used, and speech recognition technology achieves high-precision text conversion using the latest natural language processing algorithms. In the case of text input, the keyboard or touchscreen is equipped with predictive text and auto-correction functions to improve the user's input speed and accuracy. In the case of gesture input, cameras and sensors detect the user's hand movements and facial expressions with high precision, recognizing specific gestures as commands. This allows the input unit to accommodate diverse user input methods and provide an intuitive and user-friendly interface. Furthermore, the input unit records the user's input history and includes a learning function to improve the accuracy of predictions for future inputs. For example, by analyzing past input data and learning the user's input patterns and preferences, it can provide faster and more accurate input assistance. This enables the input unit to respond flexibly to user needs, significantly improving the usability of the speech support application.

[0031] The learning unit learns from past recording data and family voices based on information received by the reception unit. For example, the learning unit collects the user's past recording data and learns from that data using a generative AI. For example, the learning unit extracts the characteristics of the user's voice and generates speech based on that. The learning unit can also collect the voices of family members and learn from that data using a generative AI. For example, the learning unit collects the user's past recording data and inputs it into the generative AI. The generative AI extracts the characteristics of the user's voice and generates speech based on that. Family voices are similarly collected and input into the generative AI for learning. To efficiently process this data, the learning unit utilizes a high-performance database and parallel processing technology. For example, when extracting the characteristics of the user's voice, it performs speech waveform analysis and spectral analysis to analyze speech features such as tone, pitch, accent, and rhythm in detail. This allows the generative AI to build a model that faithfully reproduces the individuality of the user's voice. Furthermore, the learning unit uses similar methods when learning from family voices to analyze the characteristics of family voices in detail. This allows users to generate voices that mimic the voices of their family members, resulting in more natural and friendly conversations. The learning unit continuously performs these learning processes, updating the model whenever new data is added. For example, each time a user provides new recording data, the generating AI learns from that data and improves the accuracy of voice generation. The learning unit also includes an evaluation system to collect user feedback and assess the quality of the generated voices. This enables the learning unit to achieve high-quality voice generation that meets user needs, maximizing the effectiveness of the speech assistance application.

[0032] The voice generation unit generates speech that reflects individuality based on data learned by the learning unit. For example, the voice generation unit uses a generation AI to generate speech based on the characteristics of the user's voice. For example, the generation AI generates speech that reflects the tone, pitch, and accent of the user's voice. The voice generation unit can also use the generation AI to generate speech based on the characteristics of family members' voices. For example, the generation AI generates speech that reflects the tone, pitch, and accent of the user's voice. Family members' voices are also generated using the generation AI in a similar manner. The voice generation unit uses advanced speech synthesis technology to faithfully reproduce the characteristics of the user's voice based on the model learned by the generation AI. For example, the voice generation unit can utilize a deep learning-based speech synthesis model to reproduce the subtle nuances and emotions of the user's voice. As a result, the generated speech has a natural and human-like sound, providing a realistic experience as if the user were actually speaking. The voice generation unit also has a feedback loop to evaluate the quality of the generated speech and make adjustments as needed. For example, if the generated speech does not meet the user's expectations, the voice generation unit receives the feedback and readjusts the generation AI model. This allows the voice generation unit to consistently provide high-quality speech. Furthermore, the voice generation unit can support multiple speech styles and emotional expressions. For example, if a user wants to express a specific emotion, the voice generation unit adjusts the tone and pitch accordingly to produce the appropriate speech. This enables users to speak flexibly in a variety of situations. By integrating these functions and generating speech that best reflects the user's individuality, the voice generation unit can enhance the effectiveness of speech support applications.

[0033] The output unit outputs the sound generated by the sound generation unit. The output unit outputs sound using, for example, a speaker. The output unit can also output sound using headphones or earphones. Furthermore, the output unit can output sound using AR glasses. For example, the output unit outputs sound using a speaker. It can also output sound using headphones or earphones. It can also output sound using AR glasses. The output unit comprehensively manages these diverse output methods and provides optimal sound output according to the user's needs. For example, when using a speaker, the output unit automatically adjusts the volume and sound quality to provide clear and easy-to-hear sound. When using headphones or earphones, the output unit provides the optimal volume and sound quality for the user's ears, allowing them to enjoy high-quality sound while maintaining privacy. When using AR glasses, the output unit integrates sound and visual information to provide the user with a richer experience. For example, by outputting sound using the built-in speaker of the AR glasses and simultaneously displaying visual guides and information, the user can obtain information from both sound and sight. Furthermore, the output unit can automatically switch output methods according to the user's environment and situation. For example, the system can use speakers when the user is in a quiet environment and headphones or earphones when the user is in a noisy environment. Furthermore, the output unit can collect user feedback and continuously improve the quality and settings of the output audio. This allows the output unit to provide the user with optimal audio output, maximizing the effectiveness of the speech assistance application.

[0034] The recognition unit can recognize the other person's facial expressions and gestures. For example, the recognition unit can recognize the other person's facial expressions using a camera. The recognition unit can also recognize the other person's gestures using sensors. For example, the recognition unit can recognize the other person's facial expressions using a camera and detect changes in facial expressions. It can also recognize the other person's gestures using sensors and detect hand movements and head movements. Furthermore, the recognition unit can analyze the other person's facial expressions and gestures using AI. For example, the recognition unit can input image data acquired by the camera into the AI ​​and have the AI ​​perform the analysis of facial expressions and gestures. This allows for enhanced face-to-face communication by recognizing the other person's facial expressions and gestures.

[0035] The AR display unit can perform AR displays based on information recognized by the recognition unit. For example, the AR display unit can perform AR displays based on the facial expressions and gestures of the other person recognized by the recognition unit. The AR display unit can also perform AR displays based on the facial expressions and gestures of the other person recognized by the recognition unit. Furthermore, the AR display unit can use AR glasses to display the other person's facial expressions and gestures in real time. For example, the AR display unit displays on the AR glasses based on the facial expressions and gestures of the other person recognized by the recognition unit. This enhances face-to-face communication through AR displays. Some or all of the above processing in the AR display unit may be performed using, for example, a generative AI, or without a generative AI. For example, the AR display unit can input data on the other person's facial expressions and gestures recognized by the recognition unit into a generative AI, and have the generative AI generate the content of the AR display.

[0036] The reception desk can analyze a user's past input history and select the optimal input method. For example, the reception desk can store the user's past input history in a database and analyze that data using AI. The reception desk can also predict and suggest an input method to be used during a specific time period based on the user's past input history. For example, the reception desk predicts and suggests an input method to be used during a specific time period based on the user's past input history. Furthermore, the reception desk can customize the input method by referring to the content the user has entered in the past. For example, the reception desk selects the optimal input method based on the user's past input history. In this way, the reception desk can provide the optimal input method by analyzing the user's past input history. Some or all of the above processes in the reception desk may be performed using AI, or not. For example, the reception desk can input the user's past input history into AI and have AI select the optimal input method.

[0037] The reception unit can filter input based on the user's current situation and areas of interest. For example, the reception unit can prioritize displaying relevant input options based on the user's current situation. The reception unit can also filter input content based on the user's areas of interest. For example, the reception unit filters input content based on the user's areas of interest. Furthermore, the reception unit can suggest the optimal input method considering the user's current activity. For example, the reception unit suggests the optimal input method considering the user's current activity. This allows for the provision of highly relevant input by filtering based on the user's current situation and areas of interest. Some or all of the above processing in the reception unit may be performed using AI, or not. For example, the reception unit can input data on the user's current situation and areas of interest into the AI ​​and have the AI ​​perform the filtering.

[0038] The reception unit can prioritize receiving inputs that are highly relevant, taking into account the user's geographical location information. For example, if the user is in a specific location, the reception unit will prioritize receiving inputs related to that location. The reception unit can also filter relevant information based on the user's current location. For example, the reception unit will filter relevant information based on the user's current location. Furthermore, the reception unit can suggest the optimal input method, taking into account the user's geographical location information. For example, the reception unit will suggest the optimal input method, taking into account the user's geographical location information. This allows for the priority of receiving inputs that are highly relevant by considering the user's geographical location information. Some or all of the above processing in the reception unit may be performed using AI, or not. For example, the reception unit can input the user's geographical location information into AI and have AI determine the priority of highly relevant inputs.

[0039] The reception unit can analyze the user's social media activity when receiving input and accept relevant input. For example, the reception unit can prioritize accepting relevant topics from the user's social media activity. The reception unit can also analyze the content of the user's social media posts and suggest the optimal input method. Furthermore, the reception unit can customize the input content based on the user's social media activity. This allows the reception unit to provide relevant input by analyzing the user's social media activity. Some or all of the above processing in the reception unit may be performed using AI, or not. For example, the reception unit can input data on the user's social media activity into AI and have AI select relevant inputs.

[0040] The learning unit can optimize its learning algorithm by referring to past learning data during the learning process. For example, the learning unit can store past learning data in a database and analyze that data using AI. The learning unit can also select the optimal learning algorithm based on past learning data. For example, the learning unit selects the optimal learning algorithm based on past learning data. Furthermore, the learning unit can select an algorithm from past learning data that improves learning efficiency. For example, the learning unit selects an algorithm that improves learning efficiency based on past learning data. This allows the learning algorithm to be optimized by referring to past learning data. Some or all of the above processes in the learning unit may be performed using AI, or not using AI. For example, the learning unit can input past learning data into AI and have AI perform the optimization of the learning algorithm.

[0041] The learning unit can adjust the timing of learning based on the user's lifestyle patterns. For example, the learning unit can store the user's lifestyle patterns in a database and analyze that data using AI. The learning unit can also suggest the optimal learning timing based on the user's lifestyle patterns. Furthermore, the learning unit can adjust the timing of learning considering the user's daily rhythm. For example, the learning unit can adjust the timing of learning considering the user's daily rhythm. This allows for efficient learning by adjusting the timing of learning based on the user's lifestyle patterns. Some or all of the above processes in the learning unit may be performed using AI, or not. For example, the learning unit can input data on the user's lifestyle patterns into the AI ​​and have the AI ​​perform the adjustment of the learning timing.

[0042] The learning unit can weight the training data while considering the user's geographical location information. For example, if the user is in a specific location, the learning unit will give more weight to the training data related to that location. The learning unit can also weight the training data based on the user's geographical location information. For example, the learning unit will weight the training data based on the user's geographical location information. Furthermore, the learning unit can determine the priority of the training data based on the user's current location. For example, the learning unit will determine the priority of the training data based on the user's current location. This allows the training data to be weighted by considering the user's geographical location information. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input the user's geographical location information into AI and have AI perform the weighting of the training data.

[0043] The learning unit can analyze the user's social media activity during training and utilize relevant data for learning. For example, the learning unit can use relevant data from the user's social media activity for learning. The learning unit can also analyze the content of the user's social media posts and reflect it in the learning data. For example, the learning unit can analyze the content of the user's social media posts and reflect it in the learning data. Furthermore, the learning unit can customize the learning data by referring to the user's social media activity. For example, the learning unit can customize the learning data by referring to the user's social media activity. This allows relevant data to be used for learning by analyzing the user's social media activity. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input data on the user's social media activity into AI and have AI select relevant data.

[0044] The speech generation unit can improve the naturalness of the speech by referring to the user's past speech patterns during speech generation. For example, the speech generation unit can store the user's past speech patterns in a database and analyze that data using AI. The speech generation unit can also generate natural-sounding speech based on the user's past speech patterns. For example, the speech generation unit generates natural-sounding speech based on the user's past speech patterns. Furthermore, the speech generation unit can analyze the user's past speech data to improve the naturalness of the speech. For example, the speech generation unit analyzes the user's past speech data to improve the naturalness of the speech. In this way, the naturalness of the speech can be improved by referring to the user's past speech patterns. Some or all of the above processing in the speech generation unit may be performed using AI, for example, or without AI. For example, the speech generation unit can input data on the user's past speech patterns into AI and have AI perform the improvement of the naturalness of the speech.

[0045] The speech generation unit can incorporate specific phrasing and accents to reflect the user's personality during speech generation. For example, the speech generation unit can incorporate specific phrasing from the user's past speech data. The speech generation unit can also incorporate specific accents to reflect the user's personality. For example, the speech generation unit can incorporate specific accents to reflect the user's personality. Furthermore, the speech generation unit can analyze the user's speech patterns and generate speech that reflects their personality. For example, the speech generation unit can analyze the user's speech patterns and generate speech that reflects their personality. This allows for the generation of more distinctive speech by incorporating specific phrasing and accents to reflect the user's personality. Some or all of the above processing in the speech generation unit may be performed using AI, for example, or without AI. For example, the speech generation unit can input the user's past speech data into the AI ​​and have the AI ​​perform the incorporation of specific phrasing and accents.

[0046] The speech generation unit can incorporate regionally specific expressions by considering the user's geographical location information during speech generation. For example, if the user is in a specific region, the speech generation unit can incorporate regionally specific expressions. The speech generation unit can also incorporate regionally specific accents based on the user's geographical location information. For example, the speech generation unit can incorporate regionally specific accents based on the user's geographical location information. Furthermore, the speech generation unit can reflect regionally specific expressions in the speech based on the user's current location. For example, the speech generation unit can reflect regionally specific expressions in the speech based on the user's current location. In this way, by considering the user's geographical location information, regionally specific expressions can be incorporated. Some or all of the above processing in the speech generation unit may be performed using AI, for example, or without AI. For example, the speech generation unit can input the user's geographical location information into AI and have AI perform the incorporation of regionally specific expressions.

[0047] The voice generation unit can analyze the user's social media activity and reflect relevant topics in the voice during voice generation. For example, the voice generation unit can reflect relevant topics from the user's social media activity in the voice. The voice generation unit can also analyze the content of the user's social media posts and reflect them in voice generation. For example, the voice generation unit can analyze the content of the user's social media posts and reflect them in voice generation. Furthermore, the voice generation unit can optimize its voice generation algorithm by referring to the user's social media activity. For example, the voice generation unit can optimize its voice generation algorithm by referring to the user's social media activity. This allows the voice generation unit to reflect relevant topics in the voice by analyzing the user's social media activity. Some or all of the above processing in the voice generation unit may be performed using AI, for example, or without AI. For example, the voice generation unit can input data on the user's social media activity into AI and have AI perform the reflection of relevant topics.

[0048] The output unit can select the optimal output method by referring to the user's past output history when outputting. For example, the output unit can store the user's past output history in a database and analyze that data using AI. The output unit can also select the optimal output method based on the user's past output history. For example, the output unit selects the optimal output method based on the user's past output history. Furthermore, the output unit can analyze the user's past output data and optimize the output method. For example, the output unit analyzes the user's past output data and optimizes the output method. This allows the output unit to provide the optimal output method by referring to the user's past output history. Some or all of the above processing in the output unit may be performed using AI, or not. For example, the output unit can input data from the user's past output history into AI and have AI select the optimal output method.

[0049] The output unit can adjust the volume and tone of the audio based on the user's current situation when outputting. For example, the output unit can lower the volume if the user is in a quiet place. The output unit can also raise the volume if the user is in a noisy place. Furthermore, the output unit can adjust the tone of the audio according to the user's current situation. By adjusting the volume and tone of the audio according to the user's current situation, a more appropriate audio can be provided. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input data on the user's current situation into the AI ​​and have the AI ​​perform the adjustment of the volume and tone of the audio.

[0050] The output unit can prioritize outputting audio that is highly relevant, taking into account the user's geographical location information. For example, if the user is in a specific location, the output unit will prioritize outputting audio related to that location. The output unit can also filter relevant information based on the user's current location. For example, the output unit will filter relevant information based on the user's current location. Furthermore, the output unit can output the most suitable audio, taking into account the user's geographical location information. For example, the output unit will output the most suitable audio, taking into account the user's geographical location information. This allows for the priority output of highly relevant audio by considering the user's geographical location information. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input the user's geographical location information into AI and have AI determine the priority of highly relevant audio.

[0051] The output unit can analyze the user's social media activity and output relevant audio at the time of output. For example, the output unit can prioritize outputting relevant topics from the user's social media activity. The output unit can also analyze the content of the user's social media posts and output the most appropriate audio. Furthermore, the output unit can customize the audio content based on the user's social media activity. This allows the output unit to provide relevant audio by analyzing the user's social media activity. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input data on the user's social media activity into AI and have AI select relevant audio.

[0052] The recognition unit can optimize its recognition algorithm by referring to past recognition data during recognition. For example, the recognition unit can store past recognition data in a database and analyze that data using AI. The recognition unit can also select the optimal recognition algorithm based on past recognition data. For example, the recognition unit selects the optimal recognition algorithm based on past recognition data. Furthermore, the recognition unit can select an algorithm from past recognition data that improves recognition efficiency. For example, the recognition unit selects an algorithm that improves recognition efficiency based on past recognition data. In this way, the recognition algorithm can be optimized by referring to past recognition data. Some or all of the above processing in the recognition unit may be performed using AI, or not using AI. For example, the recognition unit can input past recognition data into AI and have AI perform the optimization of the recognition algorithm.

[0053] The recognition unit can prioritize the recognition of specific facial expressions and gestures to reflect the user's individuality during recognition. For example, the recognition unit can prioritize the recognition of specific facial expressions from the user's past facial expression data. The recognition unit can also prioritize the recognition of specific gestures to reflect the user's individuality. For example, the recognition unit can prioritize the recognition of specific gestures to reflect the user's individuality. Furthermore, the recognition unit can analyze the user's facial expressions and gestures to perform recognition that reflects their individuality. For example, the recognition unit can analyze the user's facial expressions and gestures to perform recognition that reflects their individuality. This allows for more personalized recognition by prioritizing the recognition of specific facial expressions and gestures to reflect the user's individuality. Some or all of the above processing in the recognition unit may be performed using AI, or not. For example, the recognition unit can input the user's past facial expression data into the AI ​​and have the AI ​​perform the preferential recognition of specific facial expressions and gestures.

[0054] The recognition unit can recognize region-specific facial expressions and gestures by considering the user's geographical location information during recognition. For example, if the user is in a specific region, the recognition unit will prioritize recognizing region-specific facial expressions. The recognition unit can also recognize region-specific gestures based on the user's geographical location information. For example, the recognition unit will recognize region-specific gestures based on the user's geographical location information. Furthermore, the recognition unit can recognize region-specific facial expressions and gestures based on the user's current location. For example, the recognition unit will recognize region-specific facial expressions and gestures based on the user's current location. In this way, by considering the user's geographical location information, region-specific facial expressions and gestures can be recognized. Some or all of the above processing in the recognition unit may be performed using AI, for example, or without AI. For example, the recognition unit can input the user's geographical location information into AI and have AI perform the recognition of region-specific facial expressions and gestures.

[0055] The recognition unit can analyze the user's social media activity during recognition and recognize relevant facial expressions and gestures. For example, the recognition unit can prioritize the recognition of relevant facial expressions from the user's social media activity. The recognition unit can also analyze the content of the user's social media posts and suggest the optimal recognition method. Furthermore, the recognition unit can customize the recognition of facial expressions and gestures by referring to the user's social media activity. For example, the recognition unit can customize the recognition of facial expressions and gestures by referring to the user's social media activity. This allows the recognition unit to recognize relevant facial expressions and gestures by analyzing the user's social media activity. Some or all of the above processing in the recognition unit may be performed using AI, or not. For example, the recognition unit can input data on the user's social media activity into AI and have AI perform the recognition of relevant facial expressions and gestures.

[0056] The AR display unit can optimize its display algorithm by referring to past display data during AR display. For example, the AR display unit can store past display data in a database and analyze that data using AI. The AR display unit can also select the optimal display algorithm based on past display data. For example, the AR display unit selects the optimal display algorithm based on past display data. Furthermore, the AR display unit can select an algorithm from past display data that improves display efficiency. For example, the AR display unit selects an algorithm that improves display efficiency based on past display data. In this way, the display algorithm can be optimized by referring to past display data. Some or all of the above processing in the AR display unit may be performed using AI, or not using AI. For example, the AR display unit can input past display data into AI and have AI perform the optimization of the display algorithm.

[0057] The AR display unit can prioritize the display of specific elements to reflect the user's personality during AR display. For example, the AR display unit can prioritize the display of specific elements based on the user's past display data. The AR display unit can also customize specific display elements to reflect the user's personality. For example, the AR display unit can customize specific display elements to reflect the user's personality. Furthermore, the AR display unit can analyze the user's display patterns and provide an AR display that reflects their personality. For example, the AR display unit can analyze the user's display patterns and provide an AR display that reflects their personality. This allows for a more personalized display by prioritizing the display of specific elements to reflect the user's personality. Some or all of the above processing in the AR display unit may be performed using AI, for example, or without AI. For example, the AR display unit can input the user's past display data into AI and have AI prioritize the display of specific elements.

[0058] The AR display unit can incorporate region-specific display elements by considering the user's geographical location information when displaying AR content. For example, if the user is in a specific region, the AR display unit will incorporate region-specific display elements. The AR display unit can also customize region-specific display elements based on the user's geographical location information. For example, the AR display unit will customize region-specific display elements based on the user's geographical location information. Furthermore, the AR display unit can prioritize the display of region-specific elements based on the user's current location. For example, the AR display unit will prioritize the display of region-specific elements based on the user's current location. This allows for the incorporation of region-specific display elements by considering the user's geographical location information. Some or all of the above processing in the AR display unit may be performed using AI, for example, or without AI. For example, the AR display unit can input the user's geographical location information into AI and have AI perform the incorporation of region-specific display elements.

[0059] The AR display unit can analyze the user's social media activity and incorporate relevant display elements during AR display. For example, the AR display unit can incorporate relevant display elements from the user's social media activity. The AR display unit can also analyze the content of the user's social media posts and suggest the most suitable display elements. Furthermore, the AR display unit can customize the display elements based on the user's social media activity. For example, the AR display unit customizes the display elements based on the user's social media activity. This allows the AR display unit to provide relevant display elements by analyzing the user's social media activity. Some or all of the above processing in the AR display unit may be performed using AI, or not. For example, the AR display unit can input data on the user's social media activity into AI and have AI select the relevant display elements.

[0060] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0061] The reception desk can analyze the user's past input patterns when receiving user input and suggest the optimal input method. For example, the reception desk can prioritize displaying input methods that the user has frequently used in the past. Furthermore, the reception desk can customize input methods considering the user's input speed and accuracy. In addition, based on the user's input history, the reception desk can suggest input methods suitable for specific time periods. This allows for the provision of more efficient input methods by considering the user's past input patterns.

[0062] The reception desk can analyze a user's past input history and select the optimal input method. For example, it can prioritize displaying input methods that the user has frequently used in the past. It can also customize input methods considering the user's input speed and accuracy. Furthermore, based on the user's input history, the reception desk can suggest input methods suitable for specific time periods. This allows for the provision of more efficient input methods by considering the user's past input history.

[0063] The reception system can filter input based on the user's current situation and areas of interest. For example, if the user is at work, the reception system will prioritize displaying work-related input options. Similarly, if the user is seeking information about their hobbies, the reception system can filter the input content to reflect those hobbies. Furthermore, the reception system can suggest the most suitable input method, taking into account the user's current activities. This allows for the provision of highly relevant input by filtering based on the user's current situation and areas of interest.

[0064] The reception system can prioritize inputs that are highly relevant to the user's geographical location when receiving input. For example, if the user is in a specific location, it will prioritize inputs related to that location. The reception system can also filter relevant information based on the user's current location. Furthermore, the reception system can suggest the optimal input method considering the user's geographical location. This allows for the priority of inputs that are highly relevant by considering the user's geographical location.

[0065] The reception desk can analyze the user's social media activity when receiving input and accept relevant input. For example, it can prioritize accepting input related to the user's social media activity. The reception desk can also analyze the content of the user's social media posts and suggest the most suitable input method. Furthermore, the reception desk can customize the input content based on the user's social media activity. This allows the system to provide relevant input by analyzing the user's social media activity.

[0066] The following briefly describes the processing flow for example form 1.

[0067] Step 1: The reception area receives user input. User input includes voice input, text input, and gesture input. For example, voice input is received via a microphone and converted to text using speech recognition technology. Text input is received via a keyboard or touchscreen, and gesture input is recognized by detecting the user's movements using cameras or sensors. Step 2: The learning unit learns from past recording data and family voices based on the information received by the reception unit. For example, it collects the user's past recording data and uses a generative AI to learn from it. The learning unit extracts the characteristics of the user's voice and generates speech based on that. It also collects the voices of family members and uses the generative AI to learn from them as well. Step 3: The voice generation unit generates speech that reflects the user's personality based on the data learned by the learning unit. For example, it uses a generation AI to generate speech that reflects the user's voice tone, pitch, accent, etc. It can also generate speech based on the characteristics of family members' voices. Step 4: The output unit outputs the sound generated by the sound generation unit. For example, the sound is output using a speaker, headphones, earphones, AR glasses, etc.

[0068] (Example of form 2) The speech support application according to an embodiment of the present invention is a novel speech support application that enables people with speech disorders to converse in a voice that is closer to their own voice. This speech support application learns the user's past recording data and the voices of family members to achieve speech synthesis that reflects the user's individuality. Specifically, it uses a generation AI to learn from the user's and their family's past recording data and generates a voice that reflects their individuality. Furthermore, by adding augmented reality (AR) functionality, it enhances face-to-face communication and naturally reflects the other person's facial expressions and gestures. For example, the speech support application collects the user's past recording data, and the generation AI learns from that data. The generation AI extracts the characteristics of the user's voice and generates speech based on that. In addition, the speech support application uses AR functionality to recognize the facial expressions and gestures of the person it is talking to in real time and provides feedback to the user. As a result, the user can have a natural conversation in a voice that reflects their individuality, improving the quality of communication. Moreover, the speech support application can be used on smartphones and AR glasses, improving the user's quality of life and increasing opportunities for social participation. This allows speech support apps to enable natural conversations using voices that reflect the user's individuality, thereby improving the quality of communication.

[0069] The speech support application according to this embodiment comprises a reception unit, a learning unit, a speech generation unit, and an output unit. The reception unit receives user input. User input includes, but is not limited to, voice input, text input, and gesture input. For example, the reception unit receives voice input via a microphone. The reception unit can also receive text input via a keyboard or touchscreen. Furthermore, the reception unit can also receive gesture input via a camera or sensor. For example, the reception unit receives voice input via a microphone and converts it to text using speech recognition technology. Text input can be entered directly using a keyboard or touchscreen. Gesture input is recognized by detecting the user's movements using a camera or sensor. The learning unit learns past recording data and family voices based on the information received by the reception unit. For example, the learning unit collects the user's past recording data and learns that data using a generation AI. For example, the learning unit extracts the characteristics of the user's voice and generates speech based on them. The learning unit can also collect the voices of family members and learn that data using a generation AI. For example, the learning unit collects the user's past recorded data and inputs it into the generating AI. The generating AI extracts the characteristics of the user's voice and generates speech based on that. Family members' voices are also collected in a similar manner and input into the generating AI for training. The speech generation unit generates speech that reflects individuality based on the data learned by the learning unit. For example, the speech generation unit uses the generating AI to generate speech based on the characteristics of the user's voice. For example, the speech generation unit generates speech that reflects the tone, pitch, and accent of the user's voice using the generating AI. The speech generation unit can also generate speech based on the characteristics of family members' voices using the generating AI. For example, the speech generation unit generates speech that reflects the tone, pitch, and accent of the user's voice using the generating AI. Family members' voices are also generated using the generating AI in a similar manner. The output unit outputs the speech generated by the speech generation unit. For example, the output unit outputs speech using a speaker. For example, the output unit outputs speech using a speaker. The output unit can also output speech using headphones or earphones. Furthermore, the output unit can also output speech using AR glasses.For example, the output unit outputs sound using a speaker. It can also output sound using headphones or earphones. It can also output sound using AR glasses. As a result, the speech support application according to the embodiment can achieve natural dialogue with a voice that reflects the user's individuality, thereby improving the quality of communication.

[0070] The reception unit receives user input. User input includes, but is not limited to, voice input, text input, and gesture input. For example, the reception unit can receive voice input via a microphone. It can also receive text input via a keyboard or touchscreen. Furthermore, it can receive gesture input via a camera or sensor. For example, the reception unit can receive voice input via a microphone and convert it to text using speech recognition technology. Text input can be done directly using a keyboard or touchscreen. Gesture input detects user movements using a camera or sensor and recognizes them as input. The reception unit comprehensively manages these diverse input methods and provides an interface to accurately understand the user's intent. For example, in the case of voice input, a highly sensitive microphone with noise cancellation is used, and speech recognition technology achieves high-precision text conversion using the latest natural language processing algorithms. In the case of text input, the keyboard or touchscreen is equipped with predictive text and auto-correction functions to improve the user's input speed and accuracy. In the case of gesture input, cameras and sensors detect the user's hand movements and facial expressions with high precision, recognizing specific gestures as commands. This allows the input unit to accommodate diverse user input methods and provide an intuitive and user-friendly interface. Furthermore, the input unit records the user's input history and includes a learning function to improve the accuracy of predictions for future inputs. For example, by analyzing past input data and learning the user's input patterns and preferences, it can provide faster and more accurate input assistance. This enables the input unit to respond flexibly to user needs, significantly improving the usability of the speech support application.

[0071] The learning unit learns from past recording data and family voices based on information received by the reception unit. For example, the learning unit collects the user's past recording data and learns from that data using a generative AI. For example, the learning unit extracts the characteristics of the user's voice and generates speech based on that. The learning unit can also collect the voices of family members and learn from that data using a generative AI. For example, the learning unit collects the user's past recording data and inputs it into the generative AI. The generative AI extracts the characteristics of the user's voice and generates speech based on that. Family voices are similarly collected and input into the generative AI for learning. To efficiently process this data, the learning unit utilizes a high-performance database and parallel processing technology. For example, when extracting the characteristics of the user's voice, it performs speech waveform analysis and spectral analysis to analyze speech features such as tone, pitch, accent, and rhythm in detail. This allows the generative AI to build a model that faithfully reproduces the individuality of the user's voice. Furthermore, the learning unit uses similar methods when learning from family voices to analyze the characteristics of family voices in detail. This allows users to generate voices that mimic the voices of their family members, resulting in more natural and friendly conversations. The learning unit continuously performs these learning processes, updating the model whenever new data is added. For example, each time a user provides new recording data, the generating AI learns from that data and improves the accuracy of voice generation. The learning unit also includes an evaluation system to collect user feedback and assess the quality of the generated voices. This enables the learning unit to achieve high-quality voice generation that meets user needs, maximizing the effectiveness of the speech assistance application.

[0072] The voice generation unit generates speech that reflects individuality based on data learned by the learning unit. For example, the voice generation unit uses a generation AI to generate speech based on the characteristics of the user's voice. For example, the generation AI generates speech that reflects the tone, pitch, and accent of the user's voice. The voice generation unit can also use the generation AI to generate speech based on the characteristics of family members' voices. For example, the generation AI generates speech that reflects the tone, pitch, and accent of the user's voice. Family members' voices are also generated using the generation AI in a similar manner. The voice generation unit uses advanced speech synthesis technology to faithfully reproduce the characteristics of the user's voice based on the model learned by the generation AI. For example, the voice generation unit can utilize a deep learning-based speech synthesis model to reproduce the subtle nuances and emotions of the user's voice. As a result, the generated speech has a natural and human-like sound, providing a realistic experience as if the user were actually speaking. The voice generation unit also has a feedback loop to evaluate the quality of the generated speech and make adjustments as needed. For example, if the generated speech does not meet the user's expectations, the voice generation unit receives the feedback and readjusts the generation AI model. This allows the voice generation unit to consistently provide high-quality speech. Furthermore, the voice generation unit can support multiple speech styles and emotional expressions. For example, if a user wants to express a specific emotion, the voice generation unit adjusts the tone and pitch accordingly to produce the appropriate speech. This enables users to speak flexibly in a variety of situations. By integrating these functions and generating speech that best reflects the user's individuality, the voice generation unit can enhance the effectiveness of speech support applications.

[0073] The output unit outputs the sound generated by the sound generation unit. The output unit outputs sound using, for example, a speaker. The output unit can also output sound using headphones or earphones. Furthermore, the output unit can output sound using AR glasses. For example, the output unit outputs sound using a speaker. It can also output sound using headphones or earphones. It can also output sound using AR glasses. The output unit comprehensively manages these diverse output methods and provides optimal sound output according to the user's needs. For example, when using a speaker, the output unit automatically adjusts the volume and sound quality to provide clear and easy-to-hear sound. When using headphones or earphones, the output unit provides the optimal volume and sound quality for the user's ears, allowing them to enjoy high-quality sound while maintaining privacy. When using AR glasses, the output unit integrates sound and visual information to provide the user with a richer experience. For example, by outputting sound using the built-in speaker of the AR glasses and simultaneously displaying visual guides and information, the user can obtain information from both sound and sight. Furthermore, the output unit can automatically switch output methods according to the user's environment and situation. For example, the system can use speakers when the user is in a quiet environment and headphones or earphones when the user is in a noisy environment. Furthermore, the output unit can collect user feedback and continuously improve the quality and settings of the output audio. This allows the output unit to provide the user with optimal audio output, maximizing the effectiveness of the speech assistance application.

[0074] The recognition unit can recognize the other person's facial expressions and gestures. For example, the recognition unit can recognize the other person's facial expressions using a camera. The recognition unit can also recognize the other person's gestures using sensors. For example, the recognition unit can recognize the other person's facial expressions using a camera and detect changes in facial expressions. It can also recognize the other person's gestures using sensors and detect hand movements and head movements. Furthermore, the recognition unit can analyze the other person's facial expressions and gestures using AI. For example, the recognition unit can input image data acquired by the camera into the AI ​​and have the AI ​​perform the analysis of facial expressions and gestures. This allows for enhanced face-to-face communication by recognizing the other person's facial expressions and gestures.

[0075] The AR display unit can perform AR displays based on information recognized by the recognition unit. For example, the AR display unit can perform AR displays based on the facial expressions and gestures of the other person recognized by the recognition unit. The AR display unit can also perform AR displays based on the facial expressions and gestures of the other person recognized by the recognition unit. Furthermore, the AR display unit can use AR glasses to display the other person's facial expressions and gestures in real time. For example, the AR display unit displays on the AR glasses based on the facial expressions and gestures of the other person recognized by the recognition unit. This enhances face-to-face communication through AR displays. Some or all of the above processing in the AR display unit may be performed using, for example, a generative AI, or without a generative AI. For example, the AR display unit can input data on the other person's facial expressions and gestures recognized by the recognition unit into a generative AI, and have the generative AI generate the content of the AR display.

[0076] The reception unit can estimate the user's emotions and adjust the timing of input reception based on the estimated emotions. For example, the reception unit can capture the user's facial expressions with a camera and estimate their emotions using an emotion estimation algorithm. The reception unit can also record the user's voice and estimate their emotions using voice analysis technology. For example, the reception unit can analyze the tone and speed of the user's voice and calculate an emotion score. Furthermore, the reception unit can collect the user's biometric data (heart rate and skin electrical activity) with sensors and estimate their emotions using an emotion estimation algorithm. For example, the reception unit can calculate an emotion score based on fluctuations in heart rate. This allows for more appropriate input reception by adjusting the timing of input reception according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes at the reception desk may be performed using AI, for example, or without AI. For example, the reception desk can input user image data captured by a camera into a generating AI and have the generating AI perform an estimation of the user's emotions.

[0077] The reception desk can analyze a user's past input history and select the optimal input method. For example, the reception desk can store the user's past input history in a database and analyze that data using AI. The reception desk can also predict and suggest an input method to be used during a specific time period based on the user's past input history. For example, the reception desk predicts and suggests an input method to be used during a specific time period based on the user's past input history. Furthermore, the reception desk can customize the input method by referring to the content the user has entered in the past. For example, the reception desk selects the optimal input method based on the user's past input history. In this way, the reception desk can provide the optimal input method by analyzing the user's past input history. Some or all of the above processes in the reception desk may be performed using AI, or not. For example, the reception desk can input the user's past input history into AI and have AI select the optimal input method.

[0078] The reception unit can filter input based on the user's current situation and areas of interest. For example, the reception unit can prioritize displaying relevant input options based on the user's current situation. The reception unit can also filter input content based on the user's areas of interest. For example, the reception unit filters input content based on the user's areas of interest. Furthermore, the reception unit can suggest the optimal input method considering the user's current activity. For example, the reception unit suggests the optimal input method considering the user's current activity. This allows for the provision of highly relevant input by filtering based on the user's current situation and areas of interest. Some or all of the above processing in the reception unit may be performed using AI, or not. For example, the reception unit can input data on the user's current situation and areas of interest into the AI ​​and have the AI ​​perform the filtering.

[0079] The reception unit can estimate the user's emotions and determine the priority of incoming inputs based on the estimated emotions. For example, the reception unit can capture the user's facial expressions with a camera and estimate their emotions using an emotion estimation algorithm. The reception unit can also record the user's voice and estimate their emotions using voice analysis technology. For example, the reception unit can analyze the tone and speed of the user's voice and calculate an emotion score. Furthermore, the reception unit can collect the user's biometric data (heart rate and skin electrical activity) with sensors and estimate their emotions using an emotion estimation algorithm. For example, the reception unit can calculate an emotion score based on fluctuations in heart rate. This allows important inputs to be processed preferentially by determining the priority of inputs according to the user's emotions. Emotion estimation is implemented using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes at the reception desk may be performed using AI, for example, or without AI. For example, the reception desk can input user image data captured by a camera into a generating AI and have the generating AI perform an estimation of the user's emotions.

[0080] The reception unit can prioritize receiving inputs that are highly relevant, taking into account the user's geographical location information. For example, if the user is in a specific location, the reception unit will prioritize receiving inputs related to that location. The reception unit can also filter relevant information based on the user's current location. For example, the reception unit will filter relevant information based on the user's current location. Furthermore, the reception unit can suggest the optimal input method, taking into account the user's geographical location information. For example, the reception unit will suggest the optimal input method, taking into account the user's geographical location information. This allows for the priority of receiving inputs that are highly relevant by considering the user's geographical location information. Some or all of the above processing in the reception unit may be performed using AI, or not. For example, the reception unit can input the user's geographical location information into AI and have AI determine the priority of highly relevant inputs.

[0081] The reception unit can analyze the user's social media activity when receiving input and accept relevant input. For example, the reception unit can prioritize accepting relevant topics from the user's social media activity. The reception unit can also analyze the content of the user's social media posts and suggest the optimal input method. Furthermore, the reception unit can customize the input content based on the user's social media activity. This allows the reception unit to provide relevant input by analyzing the user's social media activity. Some or all of the above processing in the reception unit may be performed using AI, or not. For example, the reception unit can input data on the user's social media activity into AI and have AI select relevant inputs.

[0082] The learning unit can estimate the user's emotions and select training data based on the estimated emotions. For example, the learning unit can capture the user's facial expressions with a camera and estimate the emotions using an emotion estimation algorithm. The learning unit can also record the user's voice and estimate emotions using voice analysis technology. For example, the learning unit can analyze the tone and speed of the user's voice and calculate an emotion score. Furthermore, the learning unit can collect the user's biometric data (heart rate and skin electrical activity) with sensors and estimate emotions using an emotion estimation algorithm. For example, the learning unit can calculate an emotion score based on fluctuations in heart rate. This allows for more appropriate learning by selecting training data according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the processing described above in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input user image data captured by a camera into a generating AI and have the generating AI perform the estimation of the user's emotions.

[0083] The learning unit can optimize its learning algorithm by referring to past learning data during the learning process. For example, the learning unit can store past learning data in a database and analyze that data using AI. The learning unit can also select the optimal learning algorithm based on past learning data. For example, the learning unit selects the optimal learning algorithm based on past learning data. Furthermore, the learning unit can select an algorithm from past learning data that improves learning efficiency. For example, the learning unit selects an algorithm that improves learning efficiency based on past learning data. This allows the learning algorithm to be optimized by referring to past learning data. Some or all of the above processes in the learning unit may be performed using AI, or not using AI. For example, the learning unit can input past learning data into AI and have AI perform the optimization of the learning algorithm.

[0084] The learning unit can adjust the timing of learning based on the user's lifestyle patterns. For example, the learning unit can store the user's lifestyle patterns in a database and analyze that data using AI. The learning unit can also suggest the optimal learning timing based on the user's lifestyle patterns. Furthermore, the learning unit can adjust the timing of learning considering the user's daily rhythm. For example, the learning unit can adjust the timing of learning considering the user's daily rhythm. This allows for efficient learning by adjusting the timing of learning based on the user's lifestyle patterns. Some or all of the above processes in the learning unit may be performed using AI, or not. For example, the learning unit can input data on the user's lifestyle patterns into the AI ​​and have the AI ​​perform the adjustment of the learning timing.

[0085] The learning unit can estimate the user's emotions and adjust the learning frequency based on the estimated emotions. For example, the learning unit can capture the user's facial expressions with a camera and estimate the emotions using an emotion estimation algorithm. The learning unit can also record the user's voice and estimate emotions using voice analysis technology. For example, the learning unit can analyze the tone and speed of the user's voice and calculate an emotion score. Furthermore, the learning unit can collect the user's biometric data (heart rate and skin electrical activity) with sensors and estimate emotions using an emotion estimation algorithm. For example, the learning unit can calculate an emotion score based on fluctuations in heart rate. This allows for efficient learning by adjusting the learning frequency according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the processing described above in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input user image data captured by a camera into a generating AI and have the generating AI perform the estimation of the user's emotions.

[0086] The learning unit can weight the training data while considering the user's geographical location information. For example, if the user is in a specific location, the learning unit will give more weight to the training data related to that location. The learning unit can also weight the training data based on the user's geographical location information. For example, the learning unit will weight the training data based on the user's geographical location information. Furthermore, the learning unit can determine the priority of the training data based on the user's current location. For example, the learning unit will determine the priority of the training data based on the user's current location. This allows the training data to be weighted by considering the user's geographical location information. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input the user's geographical location information into AI and have AI perform the weighting of the training data.

[0087] The learning unit can analyze the user's social media activity during training and utilize relevant data for learning. For example, the learning unit can use relevant data from the user's social media activity for learning. The learning unit can also analyze the content of the user's social media posts and reflect it in the learning data. For example, the learning unit can analyze the content of the user's social media posts and reflect it in the learning data. Furthermore, the learning unit can customize the learning data by referring to the user's social media activity. For example, the learning unit can customize the learning data by referring to the user's social media activity. This allows relevant data to be used for learning by analyzing the user's social media activity. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input data on the user's social media activity into AI and have AI select relevant data.

[0088] The voice generation unit can estimate the user's emotions and adjust the tone and pitch of the voice based on the estimated emotions. For example, the voice generation unit can capture the user's facial expressions with a camera and estimate their emotions using an emotion estimation algorithm. The voice generation unit can also record the user's voice and estimate their emotions using voice analysis technology. For example, the voice generation unit can analyze the tone and speed of the user's voice and calculate an emotion score. Furthermore, the voice generation unit can collect the user's biometric data (heart rate and skin electrical activity) with sensors and estimate their emotions using an emotion estimation algorithm. For example, the voice generation unit can calculate an emotion score based on fluctuations in heart rate. This allows for the generation of more natural-sounding voices by adjusting the tone and pitch of the voice according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the voice generation unit may be performed using AI, for example, or without AI. For example, the voice generation unit can input user image data captured by a camera into a generation AI and have the generation AI perform the estimation of the user's emotions.

[0089] The speech generation unit can improve the naturalness of the speech by referring to the user's past speech patterns during speech generation. For example, the speech generation unit can store the user's past speech patterns in a database and analyze that data using AI. The speech generation unit can also generate natural-sounding speech based on the user's past speech patterns. For example, the speech generation unit generates natural-sounding speech based on the user's past speech patterns. Furthermore, the speech generation unit can analyze the user's past speech data to improve the naturalness of the speech. For example, the speech generation unit analyzes the user's past speech data to improve the naturalness of the speech. In this way, the naturalness of the speech can be improved by referring to the user's past speech patterns. Some or all of the above processing in the speech generation unit may be performed using AI, for example, or without AI. For example, the speech generation unit can input data on the user's past speech patterns into AI and have AI perform the improvement of the naturalness of the speech.

[0090] The speech generation unit can incorporate specific phrasing and accents to reflect the user's personality during speech generation. For example, the speech generation unit can incorporate specific phrasing from the user's past speech data. The speech generation unit can also incorporate specific accents to reflect the user's personality. For example, the speech generation unit can incorporate specific accents to reflect the user's personality. Furthermore, the speech generation unit can analyze the user's speech patterns and generate speech that reflects their personality. For example, the speech generation unit can analyze the user's speech patterns and generate speech that reflects their personality. This allows for the generation of more distinctive speech by incorporating specific phrasing and accents to reflect the user's personality. Some or all of the above processing in the speech generation unit may be performed using AI, for example, or without AI. For example, the speech generation unit can input the user's past speech data into the AI ​​and have the AI ​​perform the incorporation of specific phrasing and accents.

[0091] The voice generation unit can estimate the user's emotions and adjust the speed of the speech based on the estimated emotions. For example, the voice generation unit can capture the user's facial expressions with a camera and estimate their emotions using an emotion estimation algorithm. The voice generation unit can also record the user's voice and estimate their emotions using speech analysis technology. For example, the voice generation unit analyzes the tone and speed of the user's voice and calculates an emotion score. Furthermore, the voice generation unit can collect the user's biometric data (heart rate and skin electrical activity) with sensors and estimate their emotions using an emotion estimation algorithm. For example, the voice generation unit calculates an emotion score based on fluctuations in heart rate. This allows for the generation of more natural-sounding speech by adjusting the speed of the speech according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the voice generation unit may be performed using AI, for example, or without AI. For example, the voice generation unit can input user image data captured by a camera into a generation AI and have the generation AI perform the estimation of the user's emotions.

[0092] The speech generation unit can incorporate regionally specific expressions by considering the user's geographical location information during speech generation. For example, if the user is in a specific region, the speech generation unit can incorporate regionally specific expressions. The speech generation unit can also incorporate regionally specific accents based on the user's geographical location information. For example, the speech generation unit can incorporate regionally specific accents based on the user's geographical location information. Furthermore, the speech generation unit can reflect regionally specific expressions in the speech based on the user's current location. For example, the speech generation unit can reflect regionally specific expressions in the speech based on the user's current location. In this way, by considering the user's geographical location information, regionally specific expressions can be incorporated. Some or all of the above processing in the speech generation unit may be performed using AI, for example, or without AI. For example, the speech generation unit can input the user's geographical location information into AI and have AI perform the incorporation of regionally specific expressions.

[0093] The voice generation unit can analyze the user's social media activity and reflect relevant topics in the voice during voice generation. For example, the voice generation unit can reflect relevant topics from the user's social media activity in the voice. The voice generation unit can also analyze the content of the user's social media posts and reflect them in voice generation. For example, the voice generation unit can analyze the content of the user's social media posts and reflect them in voice generation. Furthermore, the voice generation unit can optimize its voice generation algorithm by referring to the user's social media activity. For example, the voice generation unit can optimize its voice generation algorithm by referring to the user's social media activity. This allows the voice generation unit to reflect relevant topics in the voice by analyzing the user's social media activity. Some or all of the above processing in the voice generation unit may be performed using AI, for example, or without AI. For example, the voice generation unit can input data on the user's social media activity into AI and have AI perform the reflection of relevant topics.

[0094] The output unit can estimate the user's emotions and adjust the timing of the audio output based on the estimated emotions. For example, the output unit can capture the user's facial expressions with a camera and estimate their emotions using an emotion estimation algorithm. The output unit can also record the user's voice and estimate their emotions using voice analysis technology. For example, the output unit can analyze the tone and speed of the user's voice and calculate an emotion score. Furthermore, the output unit can collect the user's biometric data (heart rate and skin electrical activity) with sensors and estimate their emotions using an emotion estimation algorithm. For example, the output unit can calculate an emotion score based on fluctuations in heart rate. This allows for more appropriate timing of audio output by adjusting the timing according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input user image data captured by a camera into a generating AI and have the generating AI perform the estimation of the user's emotions.

[0095] The output unit can select the optimal output method by referring to the user's past output history when outputting. For example, the output unit can store the user's past output history in a database and analyze that data using AI. The output unit can also select the optimal output method based on the user's past output history. For example, the output unit selects the optimal output method based on the user's past output history. Furthermore, the output unit can analyze the user's past output data and optimize the output method. For example, the output unit analyzes the user's past output data and optimizes the output method. This allows the output unit to provide the optimal output method by referring to the user's past output history. Some or all of the above processing in the output unit may be performed using AI, or not. For example, the output unit can input data from the user's past output history into AI and have AI select the optimal output method.

[0096] The output unit can adjust the volume and tone of the audio based on the user's current situation when outputting. For example, the output unit can lower the volume if the user is in a quiet place. The output unit can also raise the volume if the user is in a noisy place. Furthermore, the output unit can adjust the tone of the audio according to the user's current situation. By adjusting the volume and tone of the audio according to the user's current situation, a more appropriate audio can be provided. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input data on the user's current situation into the AI ​​and have the AI ​​perform the adjustment of the volume and tone of the audio.

[0097] The output unit can estimate the user's emotions and determine the priority of the audio output based on the estimated emotions. For example, the output unit can capture the user's facial expressions with a camera and estimate their emotions using an emotion estimation algorithm. The output unit can also record the user's voice and estimate their emotions using voice analysis technology. For example, the output unit can analyze the tone and speed of the user's voice and calculate an emotion score. Furthermore, the output unit can collect the user's biometric data (heart rate and skin electrical activity) with sensors and estimate their emotions using an emotion estimation algorithm. For example, the output unit can calculate an emotion score based on fluctuations in heart rate. This allows for prioritizing audio output according to the user's emotions, thereby prioritizing the output of important audio. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input user image data captured by a camera into a generating AI and have the generating AI perform the estimation of the user's emotions.

[0098] The output unit can prioritize outputting audio that is highly relevant, taking into account the user's geographical location information. For example, if the user is in a specific location, the output unit will prioritize outputting audio related to that location. The output unit can also filter relevant information based on the user's current location. For example, the output unit will filter relevant information based on the user's current location. Furthermore, the output unit can output the most suitable audio, taking into account the user's geographical location information. For example, the output unit will output the most suitable audio, taking into account the user's geographical location information. This allows for the priority output of highly relevant audio by considering the user's geographical location information. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input the user's geographical location information into AI and have AI determine the priority of highly relevant audio.

[0099] The output unit can analyze the user's social media activity and output relevant audio at the time of output. For example, the output unit can prioritize outputting relevant topics from the user's social media activity. The output unit can also analyze the content of the user's social media posts and output the most appropriate audio. Furthermore, the output unit can customize the audio content based on the user's social media activity. This allows the output unit to provide relevant audio by analyzing the user's social media activity. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input data on the user's social media activity into AI and have AI select relevant audio.

[0100] The recognition unit can estimate the user's emotions and adjust the accuracy of facial and gesture recognition based on the estimated emotions. For example, the recognition unit can capture the user's facial expressions with a camera and estimate the emotions using an emotion estimation algorithm. The recognition unit can also record the user's voice and estimate emotions using voice analysis technology. For example, the recognition unit can analyze the tone and speed of the user's voice and calculate an emotion score. Furthermore, the recognition unit can collect the user's biometric data (heart rate and skin electrical activity) with sensors and estimate emotions using an emotion estimation algorithm. For example, the recognition unit can calculate an emotion score based on fluctuations in heart rate. This allows for more accurate recognition by adjusting the accuracy of facial and gesture recognition according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the recognition unit may be performed using AI, for example, or without AI. For example, the recognition unit can input user image data captured by a camera into a generating AI and have the generating AI perform the estimation of the user's emotions.

[0101] The recognition unit can optimize its recognition algorithm by referring to past recognition data during recognition. For example, the recognition unit can store past recognition data in a database and analyze that data using AI. The recognition unit can also select the optimal recognition algorithm based on past recognition data. For example, the recognition unit selects the optimal recognition algorithm based on past recognition data. Furthermore, the recognition unit can select an algorithm from past recognition data that improves recognition efficiency. For example, the recognition unit selects an algorithm that improves recognition efficiency based on past recognition data. In this way, the recognition algorithm can be optimized by referring to past recognition data. Some or all of the above processing in the recognition unit may be performed using AI, or not using AI. For example, the recognition unit can input past recognition data into AI and have AI perform the optimization of the recognition algorithm.

[0102] The recognition unit can prioritize the recognition of specific facial expressions and gestures to reflect the user's individuality during recognition. For example, the recognition unit can prioritize the recognition of specific facial expressions from the user's past facial expression data. The recognition unit can also prioritize the recognition of specific gestures to reflect the user's individuality. For example, the recognition unit can prioritize the recognition of specific gestures to reflect the user's individuality. Furthermore, the recognition unit can analyze the user's facial expressions and gestures to perform recognition that reflects their individuality. For example, the recognition unit can analyze the user's facial expressions and gestures to perform recognition that reflects their individuality. This allows for more personalized recognition by prioritizing the recognition of specific facial expressions and gestures to reflect the user's individuality. Some or all of the above processing in the recognition unit may be performed using AI, or not. For example, the recognition unit can input the user's past facial expression data into the AI ​​and have the AI ​​perform the preferential recognition of specific facial expressions and gestures.

[0103] The recognition unit can estimate the user's emotions and adjust the display method of the recognition results based on the estimated emotions. For example, the recognition unit can capture the user's facial expressions with a camera and estimate the emotions using an emotion estimation algorithm. The recognition unit can also record the user's voice and estimate emotions using voice analysis technology. For example, the recognition unit can analyze the tone and speed of the user's voice and calculate an emotion score. Furthermore, the recognition unit can collect the user's biometric data (heart rate and skin electrical activity) with sensors and estimate emotions using an emotion estimation algorithm. For example, the recognition unit can calculate an emotion score based on fluctuations in heart rate. This allows for a more appropriate display by adjusting the display method of the recognition results according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the recognition unit may be performed using AI, for example, or without AI. For example, the recognition unit can input user image data captured by a camera into a generating AI and have the generating AI perform the estimation of the user's emotions.

[0104] The recognition unit can recognize region-specific facial expressions and gestures by considering the user's geographical location information during recognition. For example, if the user is in a specific region, the recognition unit will prioritize recognizing region-specific facial expressions. The recognition unit can also recognize region-specific gestures based on the user's geographical location information. For example, the recognition unit will recognize region-specific gestures based on the user's geographical location information. Furthermore, the recognition unit can recognize region-specific facial expressions and gestures based on the user's current location. For example, the recognition unit will recognize region-specific facial expressions and gestures based on the user's current location. In this way, by considering the user's geographical location information, region-specific facial expressions and gestures can be recognized. Some or all of the above processing in the recognition unit may be performed using AI, for example, or without AI. For example, the recognition unit can input the user's geographical location information into AI and have AI perform the recognition of region-specific facial expressions and gestures.

[0105] The recognition unit can analyze the user's social media activity during recognition and recognize relevant facial expressions and gestures. For example, the recognition unit can prioritize the recognition of relevant facial expressions from the user's social media activity. The recognition unit can also analyze the content of the user's social media posts and suggest the optimal recognition method. Furthermore, the recognition unit can customize the recognition of facial expressions and gestures by referring to the user's social media activity. For example, the recognition unit can customize the recognition of facial expressions and gestures by referring to the user's social media activity. This allows the recognition unit to recognize relevant facial expressions and gestures by analyzing the user's social media activity. Some or all of the above processing in the recognition unit may be performed using AI, or not. For example, the recognition unit can input data on the user's social media activity into AI and have AI perform the recognition of relevant facial expressions and gestures.

[0106] The AR display unit can estimate the user's emotions and adjust the content of the AR display based on the estimated emotions. For example, the AR display unit can capture the user's facial expression with a camera and estimate the emotions using an emotion estimation algorithm. The AR display unit can also record the user's voice and estimate emotions using voice analysis technology. For example, the AR display unit can analyze the tone and speed of the user's voice and calculate an emotion score. Furthermore, the AR display unit can collect the user's biometric data (heart rate and skin electrical activity) with sensors and estimate emotions using an emotion estimation algorithm. For example, the AR display unit can calculate an emotion score based on fluctuations in heart rate. This allows for a more appropriate display by adjusting the content of the AR display according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the AR display unit may be performed using AI, for example, or without AI. For example, the AR display unit can input user image data captured by the camera into a generating AI and have the generating AI perform the estimation of the user's emotions.

[0107] The AR display unit can optimize its display algorithm by referring to past display data during AR display. For example, the AR display unit can store past display data in a database and analyze that data using AI. The AR display unit can also select the optimal display algorithm based on past display data. For example, the AR display unit selects the optimal display algorithm based on past display data. Furthermore, the AR display unit can select an algorithm from past display data that improves display efficiency. For example, the AR display unit selects an algorithm that improves display efficiency based on past display data. In this way, the display algorithm can be optimized by referring to past display data. Some or all of the above processing in the AR display unit may be performed using AI, or not using AI. For example, the AR display unit can input past display data into AI and have AI perform the optimization of the display algorithm.

[0108] The AR display unit can prioritize the display of specific elements to reflect the user's personality during AR display. For example, the AR display unit can prioritize the display of specific elements based on the user's past display data. The AR display unit can also customize specific display elements to reflect the user's personality. For example, the AR display unit can customize specific display elements to reflect the user's personality. Furthermore, the AR display unit can analyze the user's display patterns and provide an AR display that reflects their personality. For example, the AR display unit can analyze the user's display patterns and provide an AR display that reflects their personality. This allows for a more personalized display by prioritizing the display of specific elements to reflect the user's personality. Some or all of the above processing in the AR display unit may be performed using AI, for example, or without AI. For example, the AR display unit can input the user's past display data into AI and have AI prioritize the display of specific elements.

[0109] The AR display unit can estimate the user's emotions and determine the priority of AR displays based on the estimated emotions. For example, the AR display unit can capture the user's facial expressions with a camera and estimate the emotions using an emotion estimation algorithm. The AR display unit can also record the user's voice and estimate emotions using voice analysis technology. For example, the AR display unit can analyze the tone and speed of the user's voice and calculate an emotion score. Furthermore, the AR display unit can collect the user's biometric data (heart rate and skin electrical activity) with sensors and estimate emotions using an emotion estimation algorithm. For example, the AR display unit can calculate an emotion score based on fluctuations in heart rate. This allows the AR display unit to prioritize the display of important elements according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the AR display unit may be performed using AI, for example, or without AI. For example, the AR display unit can input user image data captured by the camera into a generating AI and have the generating AI perform the estimation of the user's emotions.

[0110] The AR display unit can incorporate region-specific display elements by considering the user's geographical location information when displaying AR content. For example, if the user is in a specific region, the AR display unit will incorporate region-specific display elements. The AR display unit can also customize region-specific display elements based on the user's geographical location information. For example, the AR display unit will customize region-specific display elements based on the user's geographical location information. Furthermore, the AR display unit can prioritize the display of region-specific elements based on the user's current location. For example, the AR display unit will prioritize the display of region-specific elements based on the user's current location. This allows for the incorporation of region-specific display elements by considering the user's geographical location information. Some or all of the above processing in the AR display unit may be performed using AI, for example, or without AI. For example, the AR display unit can input the user's geographical location information into AI and have AI perform the incorporation of region-specific display elements.

[0111] The AR display unit can analyze the user's social media activity and incorporate relevant display elements during AR display. For example, the AR display unit can incorporate relevant display elements from the user's social media activity. The AR display unit can also analyze the content of the user's social media posts and suggest the most suitable display elements. Furthermore, the AR display unit can customize the display elements based on the user's social media activity. For example, the AR display unit customizes the display elements based on the user's social media activity. This allows the AR display unit to provide relevant display elements by analyzing the user's social media activity. Some or all of the above processing in the AR display unit may be performed using AI, or not. For example, the AR display unit can input data on the user's social media activity into AI and have AI select the relevant display elements.

[0112] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0113] The reception desk can analyze the user's past input patterns when receiving user input and suggest the optimal input method. For example, the reception desk can prioritize displaying input methods that the user has frequently used in the past. Furthermore, the reception desk can customize input methods considering the user's input speed and accuracy. In addition, based on the user's input history, the reception desk can suggest input methods suitable for specific time periods. This allows for the provision of more efficient input methods by considering the user's past input patterns.

[0114] The recognition unit can estimate the user's emotions when recognizing the other person's facial expressions and gestures, and adjust the recognition accuracy based on the estimated emotions. For example, if the user is nervous, the recognition unit performs a more detailed analysis to improve recognition accuracy. Conversely, if the user is relaxed, the recognition unit adjusts the recognition accuracy to achieve more natural recognition. Furthermore, the recognition unit can also adjust the feedback method of the recognition results according to the user's emotions. This makes it possible to adjust the recognition accuracy according to the user's emotions, resulting in more accurate recognition.

[0115] The AR display unit can estimate the user's emotions based on the information recognized by the recognition unit and adjust the content of the AR display based on the estimated emotions. For example, if the user is excited, the AR display unit will simplify the display content to avoid information overload. Conversely, if the user is calm, the AR display unit can display detailed information. Furthermore, the AR display unit can also adjust the display color and font size according to the user's emotions. This makes it possible to adjust the AR display according to the user's emotions, enabling the provision of more appropriate information.

[0116] The reception desk can estimate the user's emotions and adjust the timing of input acceptance based on those estimates. For example, if the user is anxious, the reception desk will delay input acceptance and wait until the user calms down. Conversely, if the user is relaxed, the reception desk can speed up input acceptance. Furthermore, the reception desk can suggest input methods according to the user's emotions. This allows for adjustment of input acceptance timing according to the user's emotions, enabling more appropriate input to be received.

[0117] The reception desk can analyze a user's past input history and select the optimal input method. For example, it can prioritize displaying input methods that the user has frequently used in the past. It can also customize input methods considering the user's input speed and accuracy. Furthermore, based on the user's input history, the reception desk can suggest input methods suitable for specific time periods. This allows for the provision of more efficient input methods by considering the user's past input history.

[0118] The reception system can filter input based on the user's current situation and areas of interest. For example, if the user is at work, the reception system will prioritize displaying work-related input options. Similarly, if the user is seeking information about their hobbies, the reception system can filter the input content to reflect those hobbies. Furthermore, the reception system can suggest the most suitable input method, taking into account the user's current activities. This allows for the provision of highly relevant input by filtering based on the user's current situation and areas of interest.

[0119] The reception system can estimate the user's emotions and prioritize the inputs it receives based on those emotions. For example, if the user is excited, the reception system will prioritize important inputs. Conversely, if the user is relaxed, the reception system can prioritize normal inputs. Furthermore, the reception system can suggest input methods according to the user's emotions. This allows for the prioritization of inputs based on the user's emotions, ensuring that important inputs are processed preferentially.

[0120] The reception system can prioritize inputs that are highly relevant to the user's geographical location when receiving input. For example, if the user is in a specific location, it will prioritize inputs related to that location. The reception system can also filter relevant information based on the user's current location. Furthermore, the reception system can suggest the optimal input method considering the user's geographical location. This allows for the priority of inputs that are highly relevant by considering the user's geographical location.

[0121] The reception desk can analyze the user's social media activity when receiving input and accept relevant input. For example, it can prioritize accepting input related to the user's social media activity. The reception desk can also analyze the content of the user's social media posts and suggest the most suitable input method. Furthermore, the reception desk can customize the input content based on the user's social media activity. This allows the system to provide relevant input by analyzing the user's social media activity.

[0122] The learning unit can estimate the user's emotions and select training data based on those estimated emotions. For example, if the user is excited, the learning unit will select data appropriate for that excited state. Similarly, if the user is relaxed, the learning unit can select data appropriate for that relaxed state. Furthermore, the learning unit can adjust the timing of learning according to the user's emotions. This enables the selection of training data in accordance with the user's emotions, resulting in more effective learning.

[0123] The following briefly describes the processing flow for example form 2.

[0124] Step 1: The reception area receives user input. User input includes voice input, text input, and gesture input. For example, voice input is received via a microphone and converted to text using speech recognition technology. Text input is received via a keyboard or touchscreen, and gesture input is recognized by detecting the user's movements using cameras or sensors. Step 2: The learning unit learns from past recording data and family voices based on the information received by the reception unit. For example, it collects the user's past recording data and uses a generative AI to learn from it. The learning unit extracts the characteristics of the user's voice and generates speech based on that. It also collects the voices of family members and uses the generative AI to learn from them as well. Step 3: The voice generation unit generates speech that reflects the user's personality based on the data learned by the learning unit. For example, it uses a generation AI to generate speech that reflects the user's voice tone, pitch, accent, etc. It can also generate speech based on the characteristics of family members' voices. Step 4: The output unit outputs the sound generated by the sound generation unit. For example, the sound is output using a speaker, headphones, earphones, AR glasses, etc.

[0125] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0126] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0127] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0128] Each of the multiple elements described above, including the reception unit, learning unit, voice generation unit, output unit, recognition unit, and AR display unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the reception unit receives user input using the microphone 38B and camera 42 of the smart device 14. The learning unit learns past recorded data and family voices using the specific processing unit 290 of the data processing unit 12. The voice generation unit generates voice using a generation AI with the specific processing unit 290 of the data processing unit 12. The output unit outputs voice using the speaker 40B of the smart device 14 or AR glasses. The recognition unit recognizes the other person's facial expressions and gestures using the camera 42 of the smart device 14. The AR display unit displays the other person's facial expressions and gestures in real time using the AR glasses of the smart device 14. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.

[0129] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0130] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0131] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0132] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0133] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0134] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0135] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0136] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0137] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0138] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0139] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0140] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0141] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0142] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0143] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0144] Each of the multiple elements described above, including the reception unit, learning unit, voice generation unit, output unit, recognition unit, and AR display unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the reception unit receives user input using the microphone 238 and camera 42 of the smart glasses 214. The learning unit learns past recorded data and family voices using the specific processing unit 290 of the data processing unit 12. The voice generation unit generates voice using a generation AI via the specific processing unit 290 of the data processing unit 12. The output unit outputs voice using the speaker 240 of the smart glasses 214 or AR glasses. The recognition unit recognizes the other person's facial expressions and gestures using the camera 42 of the smart glasses 214. The AR display unit displays the other person's facial expressions and gestures in real time using the AR glasses of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.

[0145] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0146] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0147] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0148] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0149] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0150] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0151] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0152] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0153] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0154] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0155] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0156] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0157] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0158] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0159] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0160] Each of the multiple elements described above, including the reception unit, learning unit, voice generation unit, output unit, recognition unit, and AR display unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the reception unit receives user input using the microphone 238 and camera 42 of the headset terminal 314. The learning unit learns past recorded data and family voices using the specific processing unit 290 of the data processing unit 12. The voice generation unit generates voice using a generation AI with the specific processing unit 290 of the data processing unit 12. The output unit outputs voice using the speaker 240 of the headset terminal 314 or AR glasses. The recognition unit recognizes the other person's facial expressions and gestures using the camera 42 of the headset terminal 314. The AR display unit displays the other person's facial expressions and gestures in real time using the AR glasses of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.

[0161] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0162] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0163] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0164] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0165] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0166] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0167] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0168] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0169] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0170] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0171] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0172] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0173] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0174] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0175] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0176] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0177] Each of the multiple elements described above, including the reception unit, learning unit, voice generation unit, output unit, recognition unit, and AR display unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the reception unit receives user input using the microphone 238 and camera 42 of the robot 414. The learning unit learns past recorded data and family voices using the specific processing unit 290 of the data processing unit 12. The voice generation unit generates voice using a generation AI via the specific processing unit 290 of the data processing unit 12. The output unit outputs voice using the speaker 240 and AR glasses of the robot 414. The recognition unit recognizes the other person's facial expressions and gestures using the camera 42 of the robot 414. The AR display unit displays the other person's facial expressions and gestures in real time using the AR glasses of the robot 414. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0178] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0179] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0180] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0181] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0182] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0183] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0184] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0185] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0186] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0187] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0188] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0189] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0190] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0191] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0192] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0193] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0194] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0195] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0196] (Note 1) A reception area that receives user input, Based on the information received by the reception unit, the learning unit learns from past recording data and the voices of family members, A voice generation unit that generates voices that reflect individuality based on data learned by the learning unit, The system includes an output unit that outputs the sound generated by the sound generation unit. A system characterized by the following features. (Note 2) It is equipped with a recognition unit that recognizes the other party's facial expressions and gestures. The system described in Appendix 1, characterized by the features described herein. (Note 3) The system includes an AR display unit that performs AR display based on the information recognized by the recognition unit. The system described in Appendix 2, characterized by the features described herein. (Note 4) The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of input acceptance based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned reception unit is Analyze the user's past input history and select the optimal input method. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned reception unit is When receiving input, filtering is performed based on the user's current situation and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned reception unit is It estimates the user's emotions and determines the priority of input to accept based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned reception unit is When receiving input, the system prioritizes accepting inputs that are highly relevant, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned reception unit is When receiving input, the system analyzes the user's social media activity and accepts relevant input. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned learning unit, The system estimates the user's emotions and selects training data based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned learning unit, During training, the learning algorithm is optimized by referring to past training data. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned learning unit, During learning, the timing of learning is adjusted based on the user's lifestyle patterns. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned learning unit, It estimates the user's emotions and adjusts the learning frequency based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned learning unit, During training, the training data is weighted considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned learning unit, During training, the system analyzes users' social media activity and uses relevant data for learning. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned speech generation unit, It estimates the user's emotions and adjusts the tone and pitch of the voice based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned speech generation unit, When generating speech, the system improves the naturalness of the speech by referencing the user's past speech patterns. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned speech generation unit, When generating speech, specific phrases and accents are incorporated to reflect the user's individuality. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned speech generation unit, It estimates the user's emotions and adjusts the audio speed based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned speech generation unit, When generating speech, the system takes the user's geographical location into account and incorporates region-specific expressions. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned speech generation unit, During voice generation, the system analyzes the user's social media activity and incorporates relevant topics into the audio. The system described in Appendix 1, characterized by the features described herein. (Note 22) The output unit is, It estimates the user's emotions and adjusts the timing of voice output based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The output unit is, During output, the system selects the optimal output method by referring to the user's past output history. The system described in Appendix 1, characterized by the features described herein. (Note 24) The output unit is, When outputting, the volume and tone of the audio are adjusted based on the user's current state. The system described in Appendix 1, characterized by the features described herein. (Note 25) The output unit is, It estimates the user's emotions and determines the priority of the audio output based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The output unit is, When outputting audio, the system prioritizes outputting audio that is highly relevant to the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 27) The output unit is, During output, the system analyzes the user's social media activity and outputs relevant audio. The system described in Appendix 1, characterized by the features described herein. (Note 28) The recognition unit, It estimates the user's emotions and adjusts the accuracy of facial expression and gesture recognition based on the estimated emotions. The system described in Appendix 2, characterized by the features described herein. (Note 29) The recognition unit, During recognition, the recognition algorithm is optimized by referring to past recognition data. The system described in Appendix 2, characterized by the features described herein. (Note 30) The recognition unit, During recognition, specific facial expressions and gestures are prioritized to reflect the user's individuality. The system described in Appendix 2, characterized by the features described herein. (Note 31) The recognition unit, It estimates the user's emotions and adjusts how the recognition results are displayed based on the estimated emotions. The system described in Appendix 2, characterized by the features described herein. (Note 32) The recognition unit, During recognition, the system takes into account the user's geographical location to recognize region-specific facial expressions and gestures. The system described in Appendix 2, characterized by the features described herein. (Note 33) The recognition unit, During recognition, the system analyzes the user's social media activity and recognizes relevant facial expressions and gestures. The system described in Appendix 2, characterized by the features described herein. (Note 34) The AR display unit is It estimates the user's emotions and adjusts the AR display content based on the estimated emotions. The system described in Appendix 3, characterized by the features described herein. (Note 35) The AR display unit is When displaying AR, the display algorithm is optimized by referring to past display data. The system described in Appendix 3, characterized by the features described herein. (Note 36) The AR display unit is When displaying AR content, certain display elements are prioritized to reflect the user's individuality. The system described in Appendix 3, characterized by the features described herein. (Note 37) The AR display unit is It estimates the user's emotions and determines the priority of AR displays based on the estimated user emotions. The system described in Appendix 3, characterized by the features described herein. (Note 38) The AR display unit is When displaying AR content, region-specific display elements are incorporated, taking into account the user's geographical location. The system described in Appendix 3, characterized by the features described herein. (Note 39) The AR display unit is When displaying AR content, the system analyzes the user's social media activity and incorporates relevant display elements. The system described in Appendix 3, characterized by the features described herein. [Explanation of Symbols]

[0197] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A reception area that receives user input, Based on the information received by the reception unit, the learning unit learns from past recording data and the voices of family members, A voice generation unit that generates voices that reflect individuality based on data learned by the learning unit, The system includes an output unit that outputs the sound generated by the sound generation unit. A system characterized by the following features.

2. It is equipped with a recognition unit that recognizes the other party's facial expressions and gestures. The system according to feature 1.

3. The system includes an AR display unit that performs AR display based on the information recognized by the recognition unit. The system according to feature 2.

4. The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of input acceptance based on the estimated emotions. The system according to feature 1.

5. The aforementioned reception unit is Analyze the user's past input history and select the optimal input method. The system according to feature 1.

6. The aforementioned reception unit is When receiving input, filtering is performed based on the user's current situation and areas of interest. The system according to feature 1.

7. The aforementioned reception unit is It estimates the user's emotions and determines the priority of input to accept based on the estimated user emotions. The system according to feature 1.

8. The aforementioned reception unit is When receiving input, the system prioritizes accepting inputs that are highly relevant, taking into account the user's geographical location. The system according to feature 1.

9. The aforementioned reception unit is When receiving input, the system analyzes the user's social media activity and accepts relevant input. The system according to feature 1.

10. The aforementioned learning unit, The system estimates the user's emotions and selects training data based on those estimated emotions. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A