system
The system addresses the lack of cultural context in speech translation by offering real-time translations and supplementary audio, enhancing cross-lingual communication through cultural background information and phrasing suggestions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-07
- Publication Date
- 2026-04-17
AI Technical Summary
Existing speech translation systems lack the ability to provide additional information on cultural background and phrasing, leading to potential misunderstandings in cross-lingual communication.
A system comprising a reception unit, translation unit, and output unit that translates user voice input in real-time, generates supplementary audio with cultural background information, and outputs both translations and supplementary audio to facilitate smoother communication.
Enhances cross-lingual communication by providing real-time translations with cultural context and phrasing suggestions, improving user understanding and accuracy.
Smart Images

Figure 2026066662000001_ABST
Abstract
Description
Technical Field
[0004] ,
[0006] , , , , , ,
[0005] , , ,
[0003] , , , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
[0007] The system according to this embodiment can provide additional information regarding cultural background and phrasing in speech translation. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] [[ID=2I]] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The AI assistant in the voice translation app according to an embodiment of the present invention is a system that translates the user's voice input in real time, outputs the translation result as audio, suggests related phrases and expressions, and provides additional information on cultural background and phrasing. When a user makes a voice input, the AI translates the audio in real time and outputs the translation result as audio. It also provides suggestions for related phrases and expressions, as well as additional information on cultural background and phrasing. For example, if a user inputs "Hello, what's the weather like today?", this audio is translated by the AI in real time. Next, the AI translates the audio. For example, the audio "Hello, what's the weather like today?" is translated as "Hello, how is the weather today?". This translation result is output as audio. Furthermore, the AI suggests related phrases and expressions. For example, if a response such as "It's sunny today" is possible, it suggests an expression such as "It is sunny today". Additional information on cultural background and phrasing is also provided. For example, information such as "In Japan, talking about the weather is often used as part of greetings" is provided. This system allows users to translate their speech in real time and receive suggestions for relevant phrases and expressions, as well as additional information about cultural background and phrasing. This facilitates smoother communication between different languages. The AI assistant in the voice translation app translates the user's voice input in real time and generates and outputs supplementary audio, enabling smoother communication between different languages.
[0029] The AI assistant of the voice translation application according to this embodiment comprises a reception unit, a translation unit, a generation unit, and an output unit. The reception unit receives voice input from the user. User voice input includes, but is not limited to, voice input via a microphone or recorded audio files. The reception unit can receive voice in real time, for example, via a microphone. The reception unit can also read recorded audio files. Furthermore, the reception unit can perform appropriate processing according to the format of the voice input. For example, the reception unit can automatically determine the format of the voice input and perform appropriate processing. The translation unit translates the voice received by the reception unit. Translation is performed, for example, based on the translation algorithm used and the types of languages it supports, but is not limited to these examples. For example, the translation unit can translate the voice using a machine translation algorithm. The translation unit can also support multiple languages. For example, the translation unit can support multiple languages such as English, Japanese, and French. The generation unit generates supplementary audio, which is audio that provides supplementary explanations to the audio translated by the translation unit. Supplementary explanations may be based on, for example, cultural background information or technical details, but are not limited to such examples. For example, the generation unit may generate information about cultural background as supplementary audio. The generation unit may also generate information about technical details as supplementary audio. The output unit outputs the supplementary audio generated by the generation unit along with the audio translated by the translation unit. Output may be performed by, for example, audio output through a speaker or text display, but are not limited to such examples. For example, the output unit may output audio through a speaker. The output unit may also display text. As a result, the AI assistant of the voice translation app according to the embodiment can translate the user's voice input in real time and generate and output supplementary audio, thereby facilitating smooth communication between different languages. Some or all of the above-described processes in the translation unit and generation unit may be performed using, for example, a generation AI, or not using a generation AI. For example, the translation unit may input voice input to the generation AI and have the generation AI output the translation result. The generation unit may input the translation result to the generation AI and have the generation AI output the supplementary audio.
[0030] The reception unit receives voice input from the user. This voice input includes, but is not limited to, voice input via a microphone or recorded audio files. The reception unit can, for example, receive voice in real time via a microphone. It can also read recorded audio files. Furthermore, the reception unit can perform appropriate processing depending on the format of the voice input. For example, the reception unit automatically determines the format of the voice input and performs appropriate processing. Specifically, in the case of real-time voice input via a microphone, the reception unit uses noise cancellation technology to remove background noise and obtain clear audio data. In the case of recorded audio files, the reception unit automatically determines the file format (e.g., MP3, WAV, AAC, etc.) and performs appropriate decoding. Furthermore, the reception unit also has a function to automatically detect the language of the voice input, determining which language the user is speaking. This allows the reception unit to efficiently and accurately receive the user's voice input. The reception unit can also evaluate the quality of the voice input and prompt the user to re-enter if necessary. For example, if the audio is unclear or noisy, the reception desk will display a message to the user such as "Please speak again." This ensures that the reception desk always provides high-quality audio data to the translation department.
[0031] The translation department translates audio received by the reception department. Translation is performed based on, for example, the translation algorithm used and the types of languages supported, but is not limited to these examples. For example, the translation department can translate audio using a machine translation algorithm. The translation department can also handle multiple languages. For example, the translation department can handle multiple languages such as English, Japanese, and French. Specifically, the translation department uses speech recognition technology to convert audio data into text data and inputs that text data into the translation algorithm. The translation algorithm uses a neural network-based generative AI to produce highly accurate translation results. For example, in the case of translation from English to Japanese, the translation department converts English audio into text and inputs that text into the generative AI. The generative AI generates the optimal Japanese translation considering the context and meaning. Furthermore, the translation department can handle specialized terminology and slang, and can accurately translate specific industry terms and region-specific expressions used by the user. This allows the translation department to provide high-quality translations that meet the diverse needs of users. The translation department also has a function to evaluate the quality of the translation results and make corrections as needed. For example, if the translation result is unnatural, the translation unit re-inputs the text into the generation AI to produce a more appropriate translation. This allows the translation unit to consistently provide high-quality translations and improve user satisfaction.
[0032] The generation unit generates supplementary audio, which provides explanations to the audio translated by the translation unit. These explanations may be based on, for example, cultural background information or technical details, but are not limited to these examples. For instance, the generation unit can generate supplementary audio related to cultural background. It can also generate supplementary audio related to technical details. Specifically, the generation unit analyzes the translation results and inputs the necessary supplementary information into the generation AI. The generation AI generates appropriate supplementary audio based on the context and content. For example, if the translation result is "drinking tea," the generation unit might generate supplementary audio such as, "In Japan, drinking tea is an important cultural practice for relaxation and socializing." In the case of supplementary audio related to technical details, the generation unit might generate an explanation such as, "This technology uses the latest AI algorithms and provides highly accurate results." Furthermore, the generation unit can also generate individually customized supplementary audio considering the user's profile and past usage history. For example, if a user works in a specific industry, the generation unit can provide industry-specific terminology and background information as supplementary audio. This allows the generation unit to facilitate a deeper understanding for the user and improve the quality of communication. Furthermore, the generation unit also has a function to evaluate the quality of the generated supplemental audio and make corrections as needed. This allows the generation unit to consistently provide high-quality supplemental audio and improve user satisfaction.
[0033] The output unit outputs the audio translated by the translation unit along with supplementary audio generated by the generation unit. Output is performed by methods such as audio output through a speaker or text display, but is not limited to these examples. For example, the output unit outputs audio through a speaker. The output unit can also display text. Specifically, the output unit integrates the translated audio and supplementary audio to provide consistent information to the user. In the case of audio output through a speaker, the output unit plays the audio with appropriate intervals to maintain a natural flow of speech. For example, it plays the supplementary audio after the translated audio to make it easier for the user to understand the information. In the case of text display, the output unit displays the translated text and supplementary explanations on the screen, allowing the user to visually confirm the information. Furthermore, the output unit can customize the output method according to the user's device and environment. For example, it can provide a combination of audio output and text display for users using smartphones, and a more detailed text display for desktop environments. This allows the output unit to provide information to the user in the most optimal way and improve the efficiency of communication. Furthermore, the output unit can collect user feedback and use it to improve the output method. For example, if a user provides feedback on the volume or speed of the audio output, the output unit will adjust the settings accordingly. This allows the output unit to respond flexibly to user needs and improve the user experience.
[0034] The generation unit can generate supplementary audio containing additional information about cultural background. For example, the generation unit can generate supplementary audio containing information about the customs of a particular region or country. For example, the generation unit can provide supplementary audio containing information about the cultural background of Japan. The generation unit can also generate supplementary audio containing information about historical background. For example, the generation unit can provide supplementary audio containing information about a particular event or era. This allows users to gain a deeper understanding by providing additional information about cultural background and expressions. Some or all of the above processing in the generation unit may be performed using, for example, a generation AI, or without a generation AI. For example, the generation unit can input information about cultural background into a generation AI and have the generation AI output supplementary audio.
[0035] The suggestion unit can propose expressions to respond to voice. The suggestion unit proposes expressions based, for example, on appropriate phrasing and grammatical accuracy. For example, if the user asks, "What's the weather like today?", the suggestion unit will propose an expression such as "It is sunny today." Also, if the user says, "Thank you," the suggestion unit can propose an expression such as "You're welcome." In this way, by suggesting expressions to respond to voice, the user can make an appropriate response. Some or all of the above processing in the suggestion unit may be performed using AI, for example, or without AI. For example, the suggestion unit can input voice data into AI and have the AI output an appropriate expression.
[0036] The suggestion unit can propose expressions for responses, taking into account the cultural background of the audio content. For example, the suggestion unit can propose expressions based on regionally specific expressions or cultural customs. For instance, considering the cultural background of Japan, the suggestion unit might propose an expression like "Otsukaresama desu" (Thank you for your hard work). Alternatively, considering the cultural background of the United States, the suggestion unit might propose an expression like "How are you?" This allows for more appropriate responses by considering cultural backgrounds and expressions. Some or all of the processing described above in the suggestion unit may be performed using AI, or not. For example, the suggestion unit can input audio data into an AI and have the AI output an expression that takes cultural backgrounds into account.
[0037] The reception unit can analyze the user's past voice input history and select an appropriate reception method. For example, the reception unit can save past voice input history and select the optimal reception method using an analysis algorithm. For example, the reception unit can prioritize receiving phrases that the user has frequently used in the past. The reception unit can also learn the patterns of voice input used by the user in the past and suggest the optimal reception method. Furthermore, the reception unit can predict phrases to be used at specific times based on the user's past voice input history and optimize reception. In this way, the reception unit can provide the optimal reception method by analyzing past voice input history. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input past voice input history into AI and have the AI output the optimal reception method.
[0038] The reception unit can filter voice input based on the user's current situation and areas of interest. For example, the reception unit can set filtering conditions based on the user's current situation and areas of interest. For instance, if the user is in a meeting, the reception unit will prioritize receiving business-related voice input. Similarly, if the user is traveling, the reception unit can prioritize receiving voice input related to tourist attractions or transportation information. Furthermore, if the user is studying, the reception unit can prioritize receiving education-related voice input. This allows for the provision of more relevant information by filtering voice input based on the user's situation and areas of interest. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input data about the user's situation and areas of interest into the AI and have the AI output filtering conditions.
[0039] The reception unit can prioritize receiving voice input by considering the user's geographical location information. For example, the reception unit can acquire geographical location information such as GPS data or IP address and prioritize receiving voice input that is highly relevant. For example, if the user is in a specific region, the reception unit will prioritize receiving voice input related to that region. The reception unit can also prioritize receiving voice input related to the user's travel destination if the user is traveling. Furthermore, if the reception unit is at home, the reception unit can prioritize receiving voice input within the home. In this way, by considering the user's geographical location information, the reception unit can prioritize receiving voice input that is highly relevant. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input geographical location information into AI and have the AI output highly relevant voice.
[0040] The reception unit can analyze the user's social media activity when receiving voice input and receive relevant audio. For example, the reception unit can analyze the content of social media posts and the number of likes, and receive relevant audio. For example, if the reception unit is posting about a specific topic on social media, it will prioritize receiving voice input related to that topic. The reception unit can also prioritize receiving voice input related to the time of day when the user is most active on social media. Furthermore, the reception unit can prioritize receiving voice input related to areas the user has shown interest in through their social media activity. In this way, by analyzing the user's social media activity, it is possible to prioritize receiving highly relevant voice input. Some or all of the above processing in the reception unit may be performed using AI, for example, or not. For example, the reception unit can input data about social media activity into AI and have the AI output relevant audio.
[0041] The translation unit can adjust the level of detail in the translation based on the importance of the audio. For example, the translation unit can adjust the level of detail based on the urgency and relevance of the audio content. For instance, in the case of audio input of an important business meeting, the translation unit will provide a detailed and accurate translation. It can also provide a concise and easy-to-understand translation in the case of audio input of everyday conversation. Furthermore, in the case of audio input in an emergency, the translation unit can provide a quick and to-the-point translation. By adjusting the level of detail in the translation based on the importance of the audio, it is possible to provide an appropriate translation result. Some or all of the above processing in the translation unit may be performed using AI, for example, or not. For example, the translation unit can input audio data into AI and have the AI output a translation result based on importance.
[0042] The translation unit can apply different translation algorithms depending on the category of the audio during translation. For example, the translation unit can select an appropriate translation algorithm based on the category of the audio. For instance, in the case of business-related audio input, the translation unit can use a translation algorithm that corresponds to specialized terminology. Similarly, in the case of educational audio input, the translation unit can use a translation algorithm that corresponds to educational terminology. Furthermore, in the case of medical audio input, the translation unit can use a translation algorithm that corresponds to medical terminology. This improves translation accuracy by applying the appropriate translation algorithm according to the category of the audio. Some or all of the above-described processes in the translation unit may be performed using AI, for example, or without AI. For example, the translation unit can input audio data into AI and have AI apply a category-appropriate translation algorithm.
[0043] The translation unit can determine translation priorities based on the timing of speech utterances during translation. For example, the translation unit can obtain the timing of speech utterances using timestamps or the order of utterances, and then determine the translation priority. For instance, the translation unit prioritizes translation of emergency speech input. It can also prioritize translation of speech input during important meetings. Furthermore, it can prioritize translation of everyday conversations. This allows for the prioritization of important speech by determining translation priorities based on the timing of speech utterances. Some or all of the above-described processes in the translation unit may be performed using AI, or not. For example, the translation unit can input speech data into an AI and have the AI determine the translation priority based on the timing of speech utterances.
[0044] The translation unit can adjust the order of translations based on the relevance of the audio during the translation process. For example, the translation unit adjusts the order of translations based on factors such as the degree of similarity in content or common topics. For instance, in the case of audio input from an important business meeting, the translation unit will prioritize translating the most relevant parts. Similarly, in the case of audio input related to education, the translation unit can prioritize translating the most relevant parts. Furthermore, in the case of audio input related to medicine, the translation unit can prioritize translating the most relevant parts. This allows for the prioritization of important parts by adjusting the order of translations based on the relevance of the audio. Some or all of the above processing in the translation unit may be performed using AI, for example, or not. For example, the translation unit can input audio data into AI and have the AI adjust the order of translations based on relevance.
[0045] The generation unit can adjust the level of detail of supplementary audio based on the importance of the audio during the generation of supplementary audio. For example, the generation unit can adjust the level of detail of supplementary audio based on the urgency and relevance of the audio content. For example, in the case of audio input of an important business meeting, the generation unit can provide detailed and accurate supplementary audio. The generation unit can also provide concise and easy-to-understand supplementary audio in the case of audio input of everyday conversation. Furthermore, in the case of audio input in an emergency, the generation unit can provide quick and to the point. In this way, appropriate supplementary information can be provided by adjusting the level of detail of supplementary audio based on the importance of the audio. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input audio data into AI and have the AI output supplementary audio based on importance.
[0046] The generation unit can apply different supplemental speech generation algorithms depending on the category of the speech when generating supplemental speech. For example, the generation unit can select an appropriate supplemental speech generation algorithm based on the category of the speech. For example, in the case of business-related speech input, the generation unit can use a supplemental speech generation algorithm that corresponds to technical terms. The generation unit can also use a supplemental speech generation algorithm that corresponds to educational terms in the case of education-related speech input. Furthermore, the generation unit can use a supplemental speech generation algorithm that corresponds to medical terms in the case of medical-related speech input. By applying an appropriate supplemental speech generation algorithm according to the category of the speech, the accuracy of the supplemental information is improved. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input speech data into AI and have AI apply a supplemental speech generation algorithm according to the category to the AI.
[0047] The generation unit can determine the priority of supplementary audio based on the timing of speech utterances during the generation of supplementary audio. For example, the generation unit can obtain the timing of speech utterances using timestamps or the order of utterances, and then determine the priority of supplementary audio. For example, in the case of emergency speech input, the generation unit will generate supplementary audio with the highest priority. The generation unit can also generate supplementary audio with a high priority in the case of speech input during an important meeting. Furthermore, the generation unit can generate supplementary audio with a normal priority in the case of speech input from everyday conversation. By determining the priority of supplementary audio based on the timing of speech utterances, important supplementary information can be provided preferentially. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input speech data into AI and have the AI determine the priority of supplementary audio based on the timing of speech utterances.
[0048] The generation unit can adjust the order of supplementary audio based on the relevance of the audio during the generation of supplementary audio. For example, the generation unit adjusts the order of supplementary audio based on factors such as the degree of similarity in content or common topics. For example, in the case of audio input of an important business meeting, the generation unit will prioritize generating supplementary audio from highly relevant portions. The generation unit can also prioritize generating supplementary audio from highly relevant portions in the case of educational audio input. Furthermore, the generation unit can also prioritize generating supplementary audio from highly relevant portions in the case of medical audio input. This allows for prioritizing the supplementation of important portions by adjusting the order of supplementary audio based on the relevance of the audio. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input audio data into AI and have the AI adjust the order of supplementary audio based on relevance.
[0049] The output unit can select an appropriate output method by referring to the user's past voice output history when outputting voice. For example, the output unit can save past voice output history and select the optimal output method using an analysis algorithm. For example, the output unit can prioritize the use of voice output tones that the user has preferred to use in the past. The output unit can also learn the patterns of voice output used by the user in the past and suggest the optimal output method. Furthermore, the output unit can predict the tone to be used during a specific time period from the user's past voice output history and optimize the output. In this way, the optimal voice output method can be provided by referring to past voice output history. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input past voice output history into AI and have the AI output the optimal output method.
[0050] The output unit can select an appropriate output method when outputting audio, taking into account the user's device information. For example, the output unit can acquire device information such as the device type, OS, and settings, and select the optimal output method. For example, if the user is using a smartphone, the output unit will output audio that matches the speaker characteristics of the device. The output unit can also output audio optimized for a large screen if the user is using a tablet. Furthermore, if the user is using a smartwatch, the output unit can output audio that is concise and easy to read. In this way, the optimal audio output method can be provided by taking into account the user's device information. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input device information into AI and have AI output the optimal output method.
[0051] The suggestion unit can adjust the level of detail of a suggestion based on the importance of the audio. For example, the suggestion unit can adjust the level of detail based on the urgency and relevance of the audio content. For instance, in the case of an audio input of an important business meeting, the suggestion unit will provide a detailed and accurate suggestion. It can also provide a concise and easy-to-understand suggestion in the case of an audio input of an everyday conversation. Furthermore, in the case of an emergency audio input, the suggestion unit can provide a quick and to-the-point suggestion. In this way, by adjusting the level of detail of a suggestion based on the importance of the audio, it can provide an appropriate suggestion. Some or all of the above processing in the suggestion unit may be performed using AI, for example, or not using AI. For example, the suggestion unit can input audio data into AI and have the AI output a suggestion based on importance.
[0052] The suggestion unit can apply different suggestion algorithms depending on the category of the voice during the suggestion process. For example, the suggestion unit can select an appropriate suggestion algorithm based on the category of the voice. For example, in the case of business-related voice input, the suggestion unit can use a suggestion algorithm that corresponds to specialized terminology. Furthermore, in the case of education-related voice input, the suggestion unit can use a suggestion algorithm that corresponds to educational terminology. In addition, in the case of medical-related voice input, the suggestion unit can use a suggestion algorithm that corresponds to medical terminology. This improves the accuracy of the suggestions by applying an appropriate suggestion algorithm according to the category of the voice. Some or all of the above processing in the suggestion unit may be performed using AI, for example, or without AI. For example, the suggestion unit can input voice data into AI and have AI apply a suggestion algorithm appropriate to the category.
[0053] The proposal unit can determine the priority of proposals based on the timing of speech utterances. For example, the proposal unit can obtain the timing of speech utterances using timestamps or the order of utterances, and then determine the priority of proposals. For example, in the case of emergency speech input, the proposal unit will make a proposal with the highest priority. The proposal unit can also make a proposal with a high priority in the case of speech input during an important meeting. Furthermore, the proposal unit can make a proposal with a normal priority in the case of speech input from everyday conversation. In this way, by determining the priority of proposals based on the timing of speech utterances, important proposals can be given priority. Some or all of the above processing in the proposal unit may be performed using AI, for example, or without AI. For example, the proposal unit can input speech data into AI and have the AI determine the priority of proposals based on the timing of speech utterances.
[0054] The suggestion unit can adjust the order of suggestions based on the relevance of the audio during the suggestion process. For example, the suggestion unit can adjust the order of suggestions based on factors such as the degree of similarity in the audio content or common topics. For instance, in the case of audio input of an important business meeting, the suggestion unit will prioritize suggesting the most relevant parts. Similarly, in the case of audio input related to education, the suggestion unit can prioritize suggesting the most relevant parts. Furthermore, in the case of audio input related to medicine, the suggestion unit can prioritize suggesting the most relevant parts. This allows the suggestion unit to prioritize suggesting important parts by adjusting the order of suggestions based on the relevance of the audio. Some or all of the above processing in the suggestion unit may be performed using AI, for example, or not. For example, the suggestion unit can input audio data into AI and have the AI adjust the order of suggestions based on relevance.
[0055] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0056] The AI assistant in a voice translation app not only translates the user's voice input in real time, but can also analyze the user's past translation history to provide more accurate translations. For example, if the user has frequently used a particular phrase in the past, the accuracy of the translation for that phrase will be improved. Furthermore, if the user has previously pointed out mistranslations, the translation algorithm can be adjusted based on that information. In addition, by prioritizing the translation of expressions and phrases frequently used by the user based on their past translation history, more natural translation results can be provided. This enables customized translations tailored to the individual needs of the user.
[0057] The AI assistant in a voice translation app can adjust the translation results in real time, taking into account the user's current geographical location. For example, if the user is in a specific region, it can provide translations that take into account the local dialect and unique expressions. If the user is traveling, it can also provide translations based on the culture and customs of their destination. Furthermore, if the user is on a business trip, it can prioritize translating business-related terminology and expressions to provide more appropriate results. This enables translations tailored to the user's current situation and environment.
[0058] The AI assistant in a voice translation app can analyze the user's past voice input history to select the most appropriate translation method when translating the user's voice input in real time. For example, it can prioritize translating phrases the user has frequently used in the past. It can also learn the patterns of voice input the user has used in the past and suggest the optimal translation method. Furthermore, it can predict phrases the user will use at specific times of the day based on their past voice input history and optimize the translation accordingly. In this way, by analyzing past voice input history, it can provide the most optimal translation method.
[0059] The AI assistant in a voice translation app can adjust the translation results based on the user's current situation and areas of interest when translating the user's voice input in real time. For example, if the user is in a meeting, it can prioritize translating business-related terms and expressions. If the user is traveling, it can prioritize translations related to tourist attractions and transportation information. Furthermore, if the user is studying, it can prioritize translating education-related terms and expressions to provide more appropriate results. This enables translations tailored to the user's situation and areas of interest.
[0060] The AI assistant in a voice translation app can analyze the user's social media activity and provide relevant translations in real time when translating the user's voice input. For example, if the user posts about a specific topic on social media, it can prioritize translations related to that topic. It can also prioritize translations related to times when the user is most active on social media. Furthermore, by prioritizing translations related to areas the user has shown interest in through their social media activity, it can provide more appropriate translation results. In this way, by analyzing the user's social media activity, it can provide highly relevant translations.
[0061] The AI assistant in the voice translation app can apply different translation algorithms depending on the category of the voice input when translating the user's voice input in real time. For example, for business-related voice input, it can use a translation algorithm that handles specialized terminology. Similarly, for education-related voice input, it can use a translation algorithm that handles educational terminology. Furthermore, for medical-related voice input, it can use a translation algorithm that handles medical terminology. By applying the appropriate translation algorithm according to the category of the voice input, the accuracy of the translation is improved.
[0062] The following briefly describes the processing flow for example form 1.
[0063] Step 1: The reception unit receives the user's voice input. User voice input includes, for example, voice input via microphone or recorded audio files. The reception unit can receive voice in real time via microphone and can also read recorded audio files. Furthermore, the reception unit automatically determines the format of the voice input and processes it appropriately. Step 2: The translation unit translates the audio received by the reception unit. The translation is performed based on the translation algorithm used and the types of languages supported. For example, the translation unit can use a machine translation algorithm to translate the audio and can support multiple languages. Step 3: The generation unit generates supplementary audio, which provides explanations to the audio translated by the translation unit. The supplementary explanations are based on cultural background information or technical details. For example, the generation unit generates supplementary audio containing information about cultural background or technical details. Step 4: The output unit outputs the audio translated by the translation unit along with the supplementary audio generated by the generation unit. The output is performed in various ways, such as audio output through a speaker or text display. For example, the output unit can output audio through a speaker and also display text.
[0064] (Example of form 2) The AI assistant in the voice translation app according to an embodiment of the present invention is a system that translates the user's voice input in real time, outputs the translation result as audio, suggests related phrases and expressions, and provides additional information on cultural background and phrasing. When a user makes a voice input, the AI translates the audio in real time and outputs the translation result as audio. It also provides suggestions for related phrases and expressions, as well as additional information on cultural background and phrasing. For example, if a user inputs "Hello, what's the weather like today?", this audio is translated by the AI in real time. Next, the AI translates the audio. For example, the audio "Hello, what's the weather like today?" is translated as "Hello, how is the weather today?". This translation result is output as audio. Furthermore, the AI suggests related phrases and expressions. For example, if a response such as "It's sunny today" is possible, it suggests an expression such as "It is sunny today". Additional information on cultural background and phrasing is also provided. For example, information such as "In Japan, talking about the weather is often used as part of greetings" is provided. This system allows users to translate their speech in real time and receive suggestions for relevant phrases and expressions, as well as additional information about cultural background and phrasing. This facilitates smoother communication between different languages. The AI assistant in the voice translation app translates the user's voice input in real time and generates and outputs supplementary audio, enabling smoother communication between different languages.
[0065] The AI assistant of the voice translation application according to this embodiment comprises a reception unit, a translation unit, a generation unit, and an output unit. The reception unit receives voice input from the user. User voice input includes, but is not limited to, voice input via a microphone or recorded audio files. The reception unit can receive voice in real time, for example, via a microphone. The reception unit can also read recorded audio files. Furthermore, the reception unit can perform appropriate processing according to the format of the voice input. For example, the reception unit can automatically determine the format of the voice input and perform appropriate processing. The translation unit translates the voice received by the reception unit. Translation is performed, for example, based on the translation algorithm used and the types of languages it supports, but is not limited to these examples. For example, the translation unit can translate the voice using a machine translation algorithm. The translation unit can also support multiple languages. For example, the translation unit can support multiple languages such as English, Japanese, and French. The generation unit generates supplementary audio, which is audio that provides supplementary explanations to the audio translated by the translation unit. Supplementary explanations may be based on, for example, cultural background information or technical details, but are not limited to such examples. For example, the generation unit may generate information about cultural background as supplementary audio. The generation unit may also generate information about technical details as supplementary audio. The output unit outputs the supplementary audio generated by the generation unit along with the audio translated by the translation unit. Output may be performed by, for example, audio output through a speaker or text display, but are not limited to such examples. For example, the output unit may output audio through a speaker. The output unit may also display text. As a result, the AI assistant of the voice translation app according to the embodiment can translate the user's voice input in real time and generate and output supplementary audio, thereby facilitating smooth communication between different languages. Some or all of the above-described processes in the translation unit and generation unit may be performed using, for example, a generation AI, or not using a generation AI. For example, the translation unit may input voice input to the generation AI and have the generation AI output the translation result. The generation unit may input the translation result to the generation AI and have the generation AI output the supplementary audio.
[0066] The reception unit receives voice input from the user. This voice input includes, but is not limited to, voice input via a microphone or recorded audio files. The reception unit can, for example, receive voice in real time via a microphone. It can also read recorded audio files. Furthermore, the reception unit can perform appropriate processing depending on the format of the voice input. For example, the reception unit automatically determines the format of the voice input and performs appropriate processing. Specifically, in the case of real-time voice input via a microphone, the reception unit uses noise cancellation technology to remove background noise and obtain clear audio data. In the case of recorded audio files, the reception unit automatically determines the file format (e.g., MP3, WAV, AAC, etc.) and performs appropriate decoding. Furthermore, the reception unit also has a function to automatically detect the language of the voice input, determining which language the user is speaking. This allows the reception unit to efficiently and accurately receive the user's voice input. The reception unit can also evaluate the quality of the voice input and prompt the user to re-enter if necessary. For example, if the audio is unclear or noisy, the reception desk will display a message to the user such as "Please speak again." This ensures that the reception desk always provides high-quality audio data to the translation department.
[0067] The translation department translates audio received by the reception department. Translation is performed based on, for example, the translation algorithm used and the types of languages supported, but is not limited to these examples. For example, the translation department can translate audio using a machine translation algorithm. The translation department can also handle multiple languages. For example, the translation department can handle multiple languages such as English, Japanese, and French. Specifically, the translation department uses speech recognition technology to convert audio data into text data and inputs that text data into the translation algorithm. The translation algorithm uses a neural network-based generative AI to produce highly accurate translation results. For example, in the case of translation from English to Japanese, the translation department converts English audio into text and inputs that text into the generative AI. The generative AI generates the optimal Japanese translation considering the context and meaning. Furthermore, the translation department can handle specialized terminology and slang, and can accurately translate specific industry terms and region-specific expressions used by the user. This allows the translation department to provide high-quality translations that meet the diverse needs of users. The translation department also has a function to evaluate the quality of the translation results and make corrections as needed. For example, if the translation result is unnatural, the translation unit re-inputs the text into the generation AI to produce a more appropriate translation. This allows the translation unit to consistently provide high-quality translations and improve user satisfaction.
[0068] The generation unit generates supplementary audio, which provides explanations to the audio translated by the translation unit. These explanations may be based on, for example, cultural background information or technical details, but are not limited to these examples. For instance, the generation unit can generate supplementary audio related to cultural background. It can also generate supplementary audio related to technical details. Specifically, the generation unit analyzes the translation results and inputs the necessary supplementary information into the generation AI. The generation AI generates appropriate supplementary audio based on the context and content. For example, if the translation result is "drinking tea," the generation unit might generate supplementary audio such as, "In Japan, drinking tea is an important cultural practice for relaxation and socializing." In the case of supplementary audio related to technical details, the generation unit might generate an explanation such as, "This technology uses the latest AI algorithms and provides highly accurate results." Furthermore, the generation unit can also generate individually customized supplementary audio considering the user's profile and past usage history. For example, if a user works in a specific industry, the generation unit can provide industry-specific terminology and background information as supplementary audio. This allows the generation unit to facilitate a deeper understanding for the user and improve the quality of communication. Furthermore, the generation unit also has a function to evaluate the quality of the generated supplemental audio and make corrections as needed. This allows the generation unit to consistently provide high-quality supplemental audio and improve user satisfaction.
[0069] The output unit outputs the audio translated by the translation unit along with supplementary audio generated by the generation unit. Output is performed by methods such as audio output through a speaker or text display, but is not limited to these examples. For example, the output unit outputs audio through a speaker. The output unit can also display text. Specifically, the output unit integrates the translated audio and supplementary audio to provide consistent information to the user. In the case of audio output through a speaker, the output unit plays the audio with appropriate intervals to maintain a natural flow of speech. For example, it plays the supplementary audio after the translated audio to make it easier for the user to understand the information. In the case of text display, the output unit displays the translated text and supplementary explanations on the screen, allowing the user to visually confirm the information. Furthermore, the output unit can customize the output method according to the user's device and environment. For example, it can provide a combination of audio output and text display for users using smartphones, and a more detailed text display for desktop environments. This allows the output unit to provide information to the user in the most optimal way and improve the efficiency of communication. Furthermore, the output unit can collect user feedback and use it to improve the output method. For example, if a user provides feedback on the volume or speed of the audio output, the output unit will adjust the settings accordingly. This allows the output unit to respond flexibly to user needs and improve the user experience.
[0070] The generation unit can generate supplementary audio containing additional information about cultural background. For example, the generation unit can generate supplementary audio containing information about the customs of a particular region or country. For example, the generation unit can provide supplementary audio containing information about the cultural background of Japan. The generation unit can also generate supplementary audio containing information about historical background. For example, the generation unit can provide supplementary audio containing information about a particular event or era. This allows users to gain a deeper understanding by providing additional information about cultural background and expressions. Some or all of the above processing in the generation unit may be performed using, for example, a generation AI, or without a generation AI. For example, the generation unit can input information about cultural background into a generation AI and have the generation AI output supplementary audio.
[0071] The translation unit can estimate the emotions of the user who has spoken and translate based on the estimated emotions. The translation unit can estimate the user's emotions using, for example, speech analysis technology. For example, the translation unit can estimate the user's emotions by analyzing the tone and speed of the speech. The translation unit can also estimate the user's emotions using facial recognition technology. For example, the translation unit can estimate the emotions by analyzing the user's facial expressions captured by a camera. This allows for more appropriate translation results by adjusting the translation based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the translation unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the translation unit can input speech data into a generative AI and output the emotion estimation result to the generative AI.
[0072] The generation unit can estimate the emotions of the user who has uttered the speech and generate supplementary speech based on the estimated emotions of the user. The generation unit can estimate the user's emotions using, for example, speech analysis technology. For example, the generation unit can estimate the user's emotions by analyzing the tone and speed of the speech. The generation unit can also estimate the user's emotions using facial recognition technology. For example, the generation unit can estimate the emotions by analyzing the user's facial expressions captured by a camera. This allows for the provision of more appropriate supplementary information by generating supplementary speech based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the generation unit may be performed using, for example, a generation AI, or without a generation AI. For example, the generation unit can input speech data into a generation AI and have the generation AI output the emotion estimation results.
[0073] The suggestion unit can propose expressions to respond to voice. The suggestion unit proposes expressions based, for example, on appropriate phrasing and grammatical accuracy. For example, if the user asks, "What's the weather like today?", the suggestion unit will propose an expression such as "It is sunny today." Also, if the user says, "Thank you," the suggestion unit can propose an expression such as "You're welcome." In this way, by suggesting expressions to respond to voice, the user can make an appropriate response. Some or all of the above processing in the suggestion unit may be performed using AI, for example, or without AI. For example, the suggestion unit can input voice data into AI and have the AI output an appropriate expression.
[0074] The suggestion unit can propose expressions for responses, taking into account the cultural background of the audio content. For example, the suggestion unit can propose expressions based on regionally specific expressions or cultural customs. For instance, considering the cultural background of Japan, the suggestion unit might propose an expression like "Otsukaresama desu" (Thank you for your hard work). Alternatively, considering the cultural background of the United States, the suggestion unit might propose an expression like "How are you?" This allows for more appropriate responses by considering cultural backgrounds and expressions. Some or all of the processing described above in the suggestion unit may be performed using AI, or not. For example, the suggestion unit can input audio data into an AI and have the AI output an expression that takes cultural backgrounds into account.
[0075] The proposed function can estimate the emotions of a user responding to voice and propose expressions for responding based on the estimated emotions. For example, the proposed function can estimate the user's emotions using voice analysis technology. For example, the proposed function can analyze the tone and speed of the voice to estimate the user's emotions. The proposed function can also estimate the user's emotions using facial recognition technology. For example, the proposed function can analyze the user's facial expressions captured by a camera to estimate the emotions. This enables more appropriate responses by proposing expressions for responding based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the proposed function may be performed using a generative AI, or not using a generative AI. For example, the proposed function can input voice data into a generative AI and output the emotion estimation result to the generative AI.
[0076] The reception unit can estimate the user's emotions and adjust the timing of voice input reception based on the estimated emotions. The reception unit can estimate the user's emotions using, for example, voice analysis technology. For example, the reception unit can analyze the tone and speed of the voice to estimate the user's emotions. The reception unit can also estimate the user's emotions using facial recognition technology. For example, the reception unit can analyze the user's facial expressions captured by a camera to estimate their emotions. This allows for more natural conversation by adjusting the timing of voice input reception according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the reception unit can input voice data into a generative AI and have the generative AI output the emotion estimation result.
[0077] The reception unit can analyze the user's past voice input history and select an appropriate reception method. For example, the reception unit can save past voice input history and select the optimal reception method using an analysis algorithm. For example, the reception unit can prioritize receiving phrases that the user has frequently used in the past. The reception unit can also learn the patterns of voice input used by the user in the past and suggest the optimal reception method. Furthermore, the reception unit can predict phrases to be used at specific times based on the user's past voice input history and optimize reception. In this way, the reception unit can provide the optimal reception method by analyzing past voice input history. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input past voice input history into AI and have the AI output the optimal reception method.
[0078] The reception unit can filter voice input based on the user's current situation and areas of interest. For example, the reception unit can set filtering conditions based on the user's current situation and areas of interest. For instance, if the user is in a meeting, the reception unit will prioritize receiving business-related voice input. Similarly, if the user is traveling, the reception unit can prioritize receiving voice input related to tourist attractions or transportation information. Furthermore, if the user is studying, the reception unit can prioritize receiving education-related voice input. This allows for the provision of more relevant information by filtering voice input based on the user's situation and areas of interest. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input data about the user's situation and areas of interest into the AI and have the AI output filtering conditions.
[0079] The reception unit can estimate the user's emotions and determine the priority of incoming voice inputs based on the estimated emotions. The reception unit can estimate the user's emotions using, for example, voice analysis technology. For example, the reception unit can estimate the user's emotions by analyzing the tone and speed of the voice. The reception unit can also estimate the user's emotions using facial recognition technology. For example, the reception unit can estimate the emotions by analyzing the user's facial expressions captured by a camera. This allows important voice inputs to be processed preferentially by determining the priority of voice inputs based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AIs include, but are not limited to, text generation AIs (e.g., LLMs) or multimodal generation AIs. Some or all of the above processing in the reception unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the reception unit can input voice data into a generative AI and have the generative AI output the emotion estimation results.
[0080] The reception unit can prioritize receiving voice input by considering the user's geographical location information. For example, the reception unit can acquire geographical location information such as GPS data or IP address and prioritize receiving voice input that is highly relevant. For example, if the user is in a specific region, the reception unit will prioritize receiving voice input related to that region. The reception unit can also prioritize receiving voice input related to the user's travel destination if the user is traveling. Furthermore, if the reception unit is at home, the reception unit can prioritize receiving voice input within the home. In this way, by considering the user's geographical location information, the reception unit can prioritize receiving voice input that is highly relevant. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input geographical location information into AI and have the AI output highly relevant voice.
[0081] The reception unit can analyze the user's social media activity when receiving voice input and receive relevant audio. For example, the reception unit can analyze the content of social media posts and the number of likes, and receive relevant audio. For example, if the reception unit is posting about a specific topic on social media, it will prioritize receiving voice input related to that topic. The reception unit can also prioritize receiving voice input related to the time of day when the user is most active on social media. Furthermore, the reception unit can prioritize receiving voice input related to areas the user has shown interest in through their social media activity. In this way, by analyzing the user's social media activity, it is possible to prioritize receiving highly relevant voice input. Some or all of the above processing in the reception unit may be performed using AI, for example, or not. For example, the reception unit can input data about social media activity into AI and have the AI output relevant audio.
[0082] The translation unit can estimate the user's emotions and adjust the translation's expression based on the estimated emotions. For example, the translation unit can estimate the user's emotions using speech analysis technology. For instance, it can analyze the tone and speed of the speech to estimate the user's emotions. Alternatively, the translation unit can estimate the user's emotions using facial recognition technology. For example, it can analyze the user's facial expressions captured by a camera to estimate their emotions. This allows for more appropriate translation results by adjusting the translation's expression based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the translation unit may be performed using, for example, a generative AI, or without one. For example, the translation unit can input speech data into a generative AI and output the emotion estimation result to the generative AI.
[0083] The translation unit can adjust the level of detail in the translation based on the importance of the audio. For example, the translation unit can adjust the level of detail based on the urgency and relevance of the audio content. For instance, in the case of audio input of an important business meeting, the translation unit will provide a detailed and accurate translation. It can also provide a concise and easy-to-understand translation in the case of audio input of everyday conversation. Furthermore, in the case of audio input in an emergency, the translation unit can provide a quick and to-the-point translation. By adjusting the level of detail in the translation based on the importance of the audio, it is possible to provide an appropriate translation result. Some or all of the above processing in the translation unit may be performed using AI, for example, or not. For example, the translation unit can input audio data into AI and have the AI output a translation result based on importance.
[0084] The translation unit can apply different translation algorithms depending on the category of the audio during translation. For example, the translation unit can select an appropriate translation algorithm based on the category of the audio. For instance, in the case of business-related audio input, the translation unit can use a translation algorithm that corresponds to specialized terminology. Similarly, in the case of educational audio input, the translation unit can use a translation algorithm that corresponds to educational terminology. Furthermore, in the case of medical audio input, the translation unit can use a translation algorithm that corresponds to medical terminology. This improves translation accuracy by applying the appropriate translation algorithm according to the category of the audio. Some or all of the above-described processes in the translation unit may be performed using AI, for example, or without AI. For example, the translation unit can input audio data into AI and have AI apply a category-appropriate translation algorithm.
[0085] The translation unit can estimate the user's emotions and adjust the translation length based on the estimated emotions. For example, the translation unit can estimate the user's emotions using speech analysis technology. For instance, it can analyze the tone and speed of the speech to estimate the user's emotions. Alternatively, the translation unit can estimate the user's emotions using facial recognition technology. For example, it can analyze the user's facial expressions captured by a camera to estimate their emotions. This allows for more appropriate translation results by adjusting the translation length based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the translation unit may be performed using, for example, a generative AI, or not. For example, the translation unit can input speech data into a generative AI and output the emotion estimation results to the generative AI.
[0086] The translation unit can determine translation priorities based on the timing of speech utterances during translation. For example, the translation unit can obtain the timing of speech utterances using timestamps or the order of utterances, and then determine the translation priority. For instance, the translation unit prioritizes translation of emergency speech input. It can also prioritize translation of speech input during important meetings. Furthermore, it can prioritize translation of everyday conversations. This allows for the prioritization of important speech by determining translation priorities based on the timing of speech utterances. Some or all of the above-described processes in the translation unit may be performed using AI, or not. For example, the translation unit can input speech data into an AI and have the AI determine the translation priority based on the timing of speech utterances.
[0087] The translation unit can adjust the order of translations based on the relevance of the audio during the translation process. For example, the translation unit adjusts the order of translations based on factors such as the degree of similarity in content or common topics. For instance, in the case of audio input from an important business meeting, the translation unit will prioritize translating the most relevant parts. Similarly, in the case of audio input related to education, the translation unit can prioritize translating the most relevant parts. Furthermore, in the case of audio input related to medicine, the translation unit can prioritize translating the most relevant parts. This allows for the prioritization of important parts by adjusting the order of translations based on the relevance of the audio. Some or all of the above processing in the translation unit may be performed using AI, for example, or not. For example, the translation unit can input audio data into AI and have the AI adjust the order of translations based on relevance.
[0088] The generation unit can estimate the user's emotions and adjust the way supplemental audio is expressed based on the estimated user emotions. For example, the generation unit can estimate the user's emotions using speech analysis technology. For example, the generation unit can analyze the tone and speed of the speech to estimate the user's emotions. The generation unit can also estimate the user's emotions using facial recognition technology. For example, the generation unit can analyze the user's facial expressions captured by a camera to estimate their emotions. This allows for the provision of more appropriate supplemental information by adjusting the way supplemental audio is expressed based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the generation unit may be performed using a generation AI, or not. For example, the generation unit can input speech data into a generation AI and output the emotion estimation result to the generation AI.
[0089] The generation unit can adjust the level of detail of supplementary audio based on the importance of the audio during the generation of supplementary audio. For example, the generation unit can adjust the level of detail of supplementary audio based on the urgency and relevance of the audio content. For example, in the case of audio input of an important business meeting, the generation unit can provide detailed and accurate supplementary audio. The generation unit can also provide concise and easy-to-understand supplementary audio in the case of audio input of everyday conversation. Furthermore, in the case of audio input in an emergency, the generation unit can provide quick and to the point. In this way, appropriate supplementary information can be provided by adjusting the level of detail of supplementary audio based on the importance of the audio. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input audio data into AI and have the AI output supplementary audio based on importance.
[0090] The generation unit can apply different supplemental speech generation algorithms depending on the category of the speech when generating supplemental speech. For example, the generation unit can select an appropriate supplemental speech generation algorithm based on the category of the speech. For example, in the case of business-related speech input, the generation unit can use a supplemental speech generation algorithm that corresponds to technical terms. The generation unit can also use a supplemental speech generation algorithm that corresponds to educational terms in the case of education-related speech input. Furthermore, the generation unit can use a supplemental speech generation algorithm that corresponds to medical terms in the case of medical-related speech input. By applying an appropriate supplemental speech generation algorithm according to the category of the speech, the accuracy of the supplemental information is improved. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input speech data into AI and have AI apply a supplemental speech generation algorithm according to the category to the AI.
[0091] The generation unit can estimate the user's emotions and adjust the length of the supplementary audio based on the estimated emotions. The generation unit can estimate the user's emotions using, for example, speech analysis technology. For example, the generation unit can analyze the tone and speed of the speech to estimate the user's emotions. The generation unit can also estimate the user's emotions using facial recognition technology. For example, the generation unit can analyze the user's facial expressions captured by a camera to estimate their emotions. This allows for the provision of more appropriate supplementary information by adjusting the length of the supplementary audio based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the generation unit may be performed using, for example, a generation AI, or not using a generation AI. For example, the generation unit can input speech data into a generation AI and output the emotion estimation result to the generation AI.
[0092] The generation unit can determine the priority of supplementary audio based on the timing of speech utterances during the generation of supplementary audio. For example, the generation unit can obtain the timing of speech utterances using timestamps or the order of utterances, and then determine the priority of supplementary audio. For example, in the case of emergency speech input, the generation unit will generate supplementary audio with the highest priority. The generation unit can also generate supplementary audio with a high priority in the case of speech input during an important meeting. Furthermore, the generation unit can generate supplementary audio with a normal priority in the case of speech input from everyday conversation. By determining the priority of supplementary audio based on the timing of speech utterances, important supplementary information can be provided preferentially. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input speech data into AI and have the AI determine the priority of supplementary audio based on the timing of speech utterances.
[0093] The generation unit can adjust the order of supplementary audio based on the relevance of the audio during the generation of supplementary audio. For example, the generation unit adjusts the order of supplementary audio based on factors such as the degree of similarity in content or common topics. For example, in the case of audio input of an important business meeting, the generation unit will prioritize generating supplementary audio from highly relevant portions. The generation unit can also prioritize generating supplementary audio from highly relevant portions in the case of educational audio input. Furthermore, the generation unit can also prioritize generating supplementary audio from highly relevant portions in the case of medical audio input. This allows for prioritizing the supplementation of important portions by adjusting the order of supplementary audio based on the relevance of the audio. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input audio data into AI and have the AI adjust the order of supplementary audio based on relevance.
[0094] The output unit can estimate the user's emotions and adjust the expression method of the voice output based on the estimated user emotions. The output unit can estimate the user's emotions using, for example, voice analysis technology. For example, the output unit can estimate the user's emotions by analyzing the tone and speed of the voice. The output unit can also estimate the user's emotions using facial recognition technology. For example, the output unit can estimate the emotions by analyzing the user's facial expressions captured by a camera. This allows for more appropriate voice output by adjusting the expression method of the voice output based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the output unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the output unit can input voice data into a generative AI and have the generative AI output the emotion estimation result.
[0095] The output unit can select an appropriate output method by referring to the user's past voice output history when outputting voice. For example, the output unit can save past voice output history and select the optimal output method using an analysis algorithm. For example, the output unit can prioritize the use of voice output tones that the user has preferred to use in the past. The output unit can also learn the patterns of voice output used by the user in the past and suggest the optimal output method. Furthermore, the output unit can predict the tone to be used during a specific time period from the user's past voice output history and optimize the output. In this way, the optimal voice output method can be provided by referring to past voice output history. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input past voice output history into AI and have the AI output the optimal output method.
[0096] The output unit can estimate the user's emotions and determine the priority of audio output based on the estimated user emotions. The output unit can estimate the user's emotions using, for example, speech analysis technology. For example, the output unit can estimate the user's emotions by analyzing the tone and speed of the voice. The output unit can also estimate the user's emotions using facial recognition technology. For example, the output unit can estimate the emotions by analyzing the user's facial expressions captured by a camera. This allows important audio output to be prioritized by determining the priority of audio output based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the output unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the output unit can input audio data into a generative AI and have the generative AI output the emotion estimation result.
[0097] The output unit can select an appropriate output method when outputting audio, taking into account the user's device information. For example, the output unit can acquire device information such as the device type, OS, and settings, and select the optimal output method. For example, if the user is using a smartphone, the output unit will output audio that matches the speaker characteristics of the device. The output unit can also output audio optimized for a large screen if the user is using a tablet. Furthermore, if the user is using a smartwatch, the output unit can output audio that is concise and easy to read. In this way, the optimal audio output method can be provided by taking into account the user's device information. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can input device information into AI and have AI output the optimal output method.
[0098] The proposal unit can estimate the user's emotions and adjust the way the proposal is presented based on the estimated emotions. For example, the proposal unit can estimate the user's emotions using speech analysis technology. For example, the proposal unit can analyze the tone and speed of the speech to estimate the user's emotions. The proposal unit can also estimate the user's emotions using facial recognition technology. For example, the proposal unit can analyze the user's facial expressions captured by a camera to estimate their emotions. This allows for more appropriate proposals by adjusting the way the proposal is presented based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the proposal unit may be performed using a generative AI, or not using a generative AI. For example, the proposal unit can input speech data into a generative AI and output the emotion estimation result to the generative AI.
[0099] The suggestion unit can adjust the level of detail of a suggestion based on the importance of the audio. For example, the suggestion unit can adjust the level of detail based on the urgency and relevance of the audio content. For instance, in the case of an audio input of an important business meeting, the suggestion unit will provide a detailed and accurate suggestion. It can also provide a concise and easy-to-understand suggestion in the case of an audio input of an everyday conversation. Furthermore, in the case of an emergency audio input, the suggestion unit can provide a quick and to-the-point suggestion. In this way, by adjusting the level of detail of a suggestion based on the importance of the audio, it can provide an appropriate suggestion. Some or all of the above processing in the suggestion unit may be performed using AI, for example, or not using AI. For example, the suggestion unit can input audio data into AI and have the AI output a suggestion based on importance.
[0100] The suggestion unit can apply different suggestion algorithms depending on the category of the voice during the suggestion process. For example, the suggestion unit can select an appropriate suggestion algorithm based on the category of the voice. For example, in the case of business-related voice input, the suggestion unit can use a suggestion algorithm that corresponds to specialized terminology. Furthermore, in the case of education-related voice input, the suggestion unit can use a suggestion algorithm that corresponds to educational terminology. In addition, in the case of medical-related voice input, the suggestion unit can use a suggestion algorithm that corresponds to medical terminology. This improves the accuracy of the suggestions by applying an appropriate suggestion algorithm according to the category of the voice. Some or all of the above processing in the suggestion unit may be performed using AI, for example, or without AI. For example, the suggestion unit can input voice data into AI and have AI apply a suggestion algorithm appropriate to the category.
[0101] The suggestion unit can estimate the user's emotions and determine the priority of suggestions based on the estimated user emotions. For example, the suggestion unit can estimate the user's emotions using speech analysis technology. For example, the suggestion unit can estimate the user's emotions by analyzing the tone and speed of the speech. The suggestion unit can also estimate the user's emotions using facial recognition technology. For example, the suggestion unit can estimate the emotions by analyzing the user's facial expressions captured by a camera. This allows important suggestions to be prioritized by determining the priority of suggestions based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the suggestion unit may be performed using a generative AI, or not using a generative AI. For example, the suggestion unit can input speech data into a generative AI and output the emotion estimation result to the generative AI.
[0102] The proposal unit can determine the priority of proposals based on the timing of speech utterances. For example, the proposal unit can obtain the timing of speech utterances using timestamps or the order of utterances, and then determine the priority of proposals. For example, in the case of emergency speech input, the proposal unit will make a proposal with the highest priority. The proposal unit can also make a proposal with a high priority in the case of speech input during an important meeting. Furthermore, the proposal unit can make a proposal with a normal priority in the case of speech input from everyday conversation. In this way, by determining the priority of proposals based on the timing of speech utterances, important proposals can be given priority. Some or all of the above processing in the proposal unit may be performed using AI, for example, or without AI. For example, the proposal unit can input speech data into AI and have the AI determine the priority of proposals based on the timing of speech utterances.
[0103] The suggestion unit can adjust the order of suggestions based on the relevance of the audio during the suggestion process. For example, the suggestion unit can adjust the order of suggestions based on factors such as the degree of similarity in the audio content or common topics. For instance, in the case of audio input of an important business meeting, the suggestion unit will prioritize suggesting the most relevant parts. Similarly, in the case of audio input related to education, the suggestion unit can prioritize suggesting the most relevant parts. Furthermore, in the case of audio input related to medicine, the suggestion unit can prioritize suggesting the most relevant parts. This allows the suggestion unit to prioritize suggesting important parts by adjusting the order of suggestions based on the relevance of the audio. Some or all of the above processing in the suggestion unit may be performed using AI, for example, or not. For example, the suggestion unit can input audio data into AI and have the AI adjust the order of suggestions based on relevance.
[0104] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0105] The AI assistant in a voice translation app not only translates the user's voice input in real time, but can also analyze the user's past translation history to provide more accurate translations. For example, if the user has frequently used a particular phrase in the past, the accuracy of the translation for that phrase will be improved. Furthermore, if the user has previously pointed out mistranslations, the translation algorithm can be adjusted based on that information. In addition, by prioritizing the translation of expressions and phrases frequently used by the user based on their past translation history, more natural translation results can be provided. This enables customized translations tailored to the individual needs of the user.
[0106] The AI assistant in a voice translation app can adjust the translation results in real time, taking into account the user's current geographical location. For example, if the user is in a specific region, it can provide translations that take into account the local dialect and unique expressions. If the user is traveling, it can also provide translations based on the culture and customs of their destination. Furthermore, if the user is on a business trip, it can prioritize translating business-related terminology and expressions to provide more appropriate results. This enables translations tailored to the user's current situation and environment.
[0107] The AI assistant in a voice translation app can estimate the user's emotions and adjust the translation results based on those emotions when translating the user's voice input in real time. For example, if the user is angry, the translation results will be adjusted to be more gentle, taking that emotion into consideration. If the user is happy, it can provide cheerful expressions that reflect that emotion. Furthermore, if the user is sad, it can provide gentle expressions that take that emotion into consideration, resulting in a more appropriate translation. This enables translations that are tailored to the user's emotions, leading to more natural communication.
[0108] The AI assistant in a voice translation app can analyze the user's past voice input history to select the most appropriate translation method when translating the user's voice input in real time. For example, it can prioritize translating phrases the user has frequently used in the past. It can also learn the patterns of voice input the user has used in the past and suggest the optimal translation method. Furthermore, it can predict phrases the user will use at specific times of the day based on their past voice input history and optimize the translation accordingly. In this way, by analyzing past voice input history, it can provide the most optimal translation method.
[0109] The AI assistant in a voice translation app can estimate the user's emotions and adjust the translation based on those emotions when translating the user's voice input in real time. For example, if the user is nervous, it can provide expressions that soothe that emotion. If the user is excited, it can provide lively expressions that reflect that emotion. Furthermore, if the user is calm, it can provide calm expressions that match that emotion, resulting in a more appropriate translation. This enables translation that responds to the user's emotions, leading to more natural communication.
[0110] The AI assistant in a voice translation app can adjust the translation results based on the user's current situation and areas of interest when translating the user's voice input in real time. For example, if the user is in a meeting, it can prioritize translating business-related terms and expressions. If the user is traveling, it can prioritize translations related to tourist attractions and transportation information. Furthermore, if the user is studying, it can prioritize translating education-related terms and expressions to provide more appropriate results. This enables translations tailored to the user's situation and areas of interest.
[0111] The AI assistant in a voice translation app can estimate the user's emotions and adjust the translation length based on those emotions when translating the user's voice input in real time. For example, if the user is in a hurry, it can provide a short and concise translation that takes that emotion into account. If the user is relaxed, it can provide a more detailed translation that matches that emotion. Furthermore, if the user is focused, it can provide a translation of an appropriate length that matches that emotion, resulting in a more accurate translation. This enables translation that is sensitive to the user's emotions, leading to more natural communication.
[0112] The AI assistant in a voice translation app can analyze the user's social media activity and provide relevant translations in real time when translating the user's voice input. For example, if the user posts about a specific topic on social media, it can prioritize translations related to that topic. It can also prioritize translations related to times when the user is most active on social media. Furthermore, by prioritizing translations related to areas the user has shown interest in through their social media activity, it can provide more appropriate translation results. In this way, by analyzing the user's social media activity, it can provide highly relevant translations.
[0113] The AI assistant in the voice translation app can estimate the user's emotions when translating their voice input in real time, and prioritize translations based on those emotions. For example, if the user is in an urgent situation, the translation will be given top priority, taking those emotions into consideration. Similarly, if the user is in an important meeting, the translation can be given a high priority based on those emotions. Furthermore, if the user is having a casual conversation, the translation can be given a normal priority according to those emotions, resulting in more appropriate translations. In this way, by prioritizing translations based on the user's emotions, important translations can be given priority.
[0114] The AI assistant in the voice translation app can apply different translation algorithms depending on the category of the voice input when translating the user's voice input in real time. For example, for business-related voice input, it can use a translation algorithm that handles specialized terminology. Similarly, for education-related voice input, it can use a translation algorithm that handles educational terminology. Furthermore, for medical-related voice input, it can use a translation algorithm that handles medical terminology. By applying the appropriate translation algorithm according to the category of the voice input, the accuracy of the translation is improved.
[0115] The following briefly describes the processing flow for example form 2.
[0116] Step 1: The reception unit receives the user's voice input. User voice input includes, for example, voice input via microphone or recorded audio files. The reception unit can receive voice in real time via microphone and can also read recorded audio files. Furthermore, the reception unit automatically determines the format of the voice input and processes it appropriately. Step 2: The translation unit translates the audio received by the reception unit. The translation is performed based on the translation algorithm used and the types of languages supported. For example, the translation unit can use a machine translation algorithm to translate the audio and can support multiple languages. Step 3: The generation unit generates supplementary audio, which provides explanations to the audio translated by the translation unit. The supplementary explanations are based on cultural background information or technical details. For example, the generation unit generates supplementary audio containing information about cultural background or technical details. Step 4: The output unit outputs the audio translated by the translation unit along with the supplementary audio generated by the generation unit. The output is performed in various ways, such as audio output through a speaker or text display. For example, the output unit can output audio through a speaker and also display text.
[0117] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0118] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0119] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0120] For example, the reception unit can receive voice input from the user through the microphone 38B of the smart device 14. The translation unit can translate the voice using the specific processing unit 290 of the data processing device 12. The generation unit can generate supplementary audio using the specific processing unit 290 of the data processing device 12. The output unit can output the translation result and supplementary audio through the speaker 40B of the smart device 14. The suggestion unit can suggest an appropriate expression using the specific processing unit 290 of the data processing device 12. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.
[0121] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0122] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0123] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0124] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0125] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0126] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0127] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0128] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0129] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0130] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0131] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0132] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0133] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0134] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0135] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0136] For example, the reception unit can receive the user's voice input through the microphone 238 of the smart glasses 214. The translation unit can translate the voice using the specific processing unit 290 of the data processing device 12. The generation unit can generate supplementary audio using the specific processing unit 290 of the data processing device 12. The output unit can output the translation result and supplementary audio through the speaker 240 of the smart glasses 214. The suggestion unit can suggest an appropriate expression using the specific processing unit 290 of the data processing device 12. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.
[0137] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0138] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0139] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0140] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0141] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0142] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0143] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0144] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0145] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0146] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0147] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0148] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0149] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0150] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0151] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0152] For example, the reception unit can receive user voice input through the microphone 238 of the headset terminal 314. The translation unit can translate the voice using the specific processing unit 290 of the data processing device 12. The generation unit can generate supplementary audio using the specific processing unit 290 of the data processing device 12. The output unit can output the translation result and supplementary audio through the speaker 240 of the headset terminal 314. The suggestion unit can suggest appropriate expressions using the specific processing unit 290 of the data processing device 12. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.
[0153] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0154] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0155] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0156] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0157] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0158] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0159] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0160] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0161] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0162] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0163] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0164] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0165] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0166] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0167] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0168] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0169] For example, the reception unit can receive voice input from the user through the microphone 238 of the robot 414. The translation unit can translate the voice using the specific processing unit 290 of the data processing device 12. The generation unit can generate supplementary voice using the specific processing unit 290 of the data processing device 12. The output unit can output the translation result and supplementary voice through the speaker 240 of the robot 414. The suggestion unit can suggest an appropriate expression using the specific processing unit 290 of the data processing device 12. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.
[0170] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0171] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0172] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0173] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0174] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0175] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0176] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0177] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0178] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0179] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0180] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0181] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0182] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0183] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0184] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0185] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0186] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0187] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0188] (Note 1) A reception desk that accepts voice input, A translation unit that translates the audio received by the reception unit, A generation unit generates audio that provides supplementary explanations to the audio translated by the translation unit, The system includes an output unit that outputs supplementary audio generated by the generation unit along with the audio translated by the translation unit. A system characterized by the following features. (Note 2) The generating unit is Generate supplementary audio to provide additional information about the cultural background. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned translation department, It estimates the emotions of the user who spoke the audio and translates based on the estimated emotions of the user. The system described in Appendix 1, characterized by the features described herein. (Note 4) The generating unit is It estimates the emotions of the user who uttered the speech and generates supplementary speech based on the estimated emotions of the user. The system described in Appendix 1, characterized by the features described herein. (Note 5) It also includes a proposal section that suggests expressions that respond to voice. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned proposal section is, Considering the cultural context of the audio content, we propose expressions for responding. The system described in Appendix 2, characterized by the features described herein. (Note 7) The aforementioned proposal section is, This system estimates the emotions of users responding to voice commands and proposes expressions for responses based on those estimated emotions. The system described in Appendix 2, characterized by the features described herein. (Note 8) The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of voice input acceptance based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned reception unit is The system analyzes the user's past voice input history and selects the appropriate reception method. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned reception unit is When receiving voice input, filtering is performed based on the user's current status. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned reception unit is It estimates the user's emotions and determines the priority of voice input to accept based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned reception unit is When receiving voice input, the system prioritizes receiving voice messages that are highly relevant based on the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned reception unit is When receiving voice input, the system accepts relevant voice messages based on the user's social media activity. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned translation department, It estimates the user's emotions and adjusts the translation's expression based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned translation department, During translation, adjust the level of detail based on the importance of the audio. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned translation department, During translation, different translation algorithms are applied depending on the audio category. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned translation department, It estimates the user's sentiment and adjusts the translation length based on the estimated sentiment. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned translation department, During translation, the translation priority is determined based on the timing of the spoken audio. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned translation department, During translation, the order of translations is adjusted based on the relevance of the audio. The system described in Appendix 1, characterized by the features described herein. (Note 20) The generating unit is The system estimates the user's emotions and adjusts the way supplementary audio is expressed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 21) The generating unit is When generating supplemental audio, adjust the level of detail in the supplemental audio based on the importance of the audio. The system described in Appendix 1, characterized by the features described herein. (Note 22) The generating unit is When generating supplemental audio, different supplemental audio generation algorithms are applied depending on the audio category. The system described in Appendix 1, characterized by the features described herein. (Note 23) The generating unit is It estimates the user's emotions and adjusts the length of the supplementary audio based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 24) The generating unit is When generating supplemental speech, the priority of supplemental speech is determined based on the timing of the speech utterances. The system described in Appendix 1, characterized by the features described herein. (Note 25) The generating unit is When generating supplemental audio, the order of the supplemental audio is adjusted based on the relevance of the audio. The system described in Appendix 1, characterized by the features described herein. (Note 26) The output unit is, It estimates the user's emotions and adjusts the way the voice output is expressed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 27) The output unit is, When outputting audio, the system refers to the user's past audio output history to select the appropriate output method. The system described in Appendix 1, characterized by the features described herein. (Note 28) The output unit is, It estimates the user's emotions and determines the priority of voice output based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The output unit is, When outputting audio, the system selects the appropriate output method considering the user's device information. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned proposal section is, It estimates the user's emotions and adjusts the way suggestions are presented based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned proposal section is, When making a proposal, adjust the level of detail in the proposal based on the importance of the voice. The system described in Appendix 1, characterized by the features described herein. (Note 32) The aforementioned proposal section is, When making a proposal, different proposal algorithms are applied depending on the audio category. The system described in Appendix 1, characterized by the features described herein. (Note 33) The aforementioned proposal section is, It estimates the user's emotions and determines the priority of suggestions based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 34) The aforementioned proposal section is, When making a proposal, prioritize the proposal based on the timing of the audio utterances. The system described in Appendix 1, characterized by the features described herein. (Note 35) The aforementioned proposal section is, When making suggestions, adjust the order of suggestions based on the relevance of the audio. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]
[0189] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A reception desk that accepts voice input, A translation unit that translates the audio received by the reception unit, A generation unit generates audio that provides supplementary explanations to the audio translated by the translation unit, The system includes an output unit that outputs supplementary audio generated by the generation unit along with the audio translated by the translation unit. A system characterized by the following features.
2. The generating unit is Generate supplementary audio to provide additional information about the cultural background. The system according to feature 1.
3. The aforementioned translation department, It estimates the emotions of the user who spoke the audio and translates based on the estimated emotions of the user. The system according to feature 1.
4. The generating unit is It estimates the emotions of the user who uttered the speech and generates supplementary speech based on the estimated emotions of the user. The system according to feature 1.
5. It also includes a proposal section that suggests expressions that respond to voice. The system according to feature 1.
6. The aforementioned proposal section is, Considering the cultural context of the audio content, we propose expressions for responding. The system according to claim 5, characterized in that it is the same as described in claim 5.
7. The aforementioned proposal section is, This system estimates the emotions of users responding to voice commands and proposes expressions for responses based on those estimated emotions. The system according to claim 5, characterized in that it is the same as described in claim 5.
8. The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of voice input acceptance based on the estimated emotions. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A