system

The system addresses the challenge of natural conversation with locals by using a reception, learning, and analysis unit to translate user input from social media and television data, allowing effective communication without language learning.

JP2026054902APending Publication Date: 2026-03-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-17
Publication Date
2026-03-30

AI Technical Summary

Technical Problem

Existing systems face difficulties in enabling natural conversations with local people without language learning.

Method used

A system comprising a reception unit, learning unit, and analysis unit that receives user input, learns from social media and television data, and provides translation results, utilizing natural language processing and machine learning algorithms to facilitate natural conversations.

Benefits of technology

Enables users to converse naturally with locals without language learning by providing accurate and contextually relevant translations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026054902000001_ABST
    Figure 2026054902000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to enable users to converse naturally with local people without language learning. [Solution] The system according to the embodiment comprises a reception unit, a learning unit, an analysis unit, and a provision unit. The reception unit receives user input. The learning unit learns based on data from social media or television. The analysis unit analyzes the input words. The provision unit provides the translation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, there is a problem that it is difficult to have a natural conversation with local people without language learning.

[0005] The system according to the embodiment aims to enable natural conversation with local people without language learning.

Means for Solving the Problems

[0006] The system according to the embodiment includes a reception unit, a learning unit, an analysis unit, and a provision unit. The reception unit receives user input. The learning unit learns based on data from SNS or TV. The analysis unit analyzes the input words. The provision unit provides the translation result.

Effects of the Invention

[0007] The system according to this embodiment can enable users to converse naturally with local people without language learning. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The AI ​​translation system according to an embodiment of the present invention is a system that provides a translation service that allows users to converse naturally with locals without having to learn the language. This AI translation system works by having the user input the words they want to translate, and the AI ​​performs a natural translation based on learning data of words used on social media and television. This learning data is regularly updated, allowing the system to handle the latest words and expressions. For example, in Japan, unique business phrases such as "Let's all work together" and slang terms like "Haoi" are constantly emerging. The AI ​​learns these words and expressions and translates them in a way that foreigners can naturally understand. For example, the user inputs a business phrase like "Let's all work together" or slang terms like "Haoi." This information is input to the AI. Next, the AI ​​analyzes the input words and performs a natural translation based on learning data of words used on social media and television. By regularly updating this learning data, the AI ​​can handle the latest words and expressions. For example, it translates the phrase "Let's all work together" into an expression that is easy for foreigners to understand. This mechanism allows users to converse naturally with locals without having to learn the language. For example, AI can naturally translate words and expressions that are difficult for foreigners to fully understand, even after extensive Japanese study, enabling them to communicate smoothly with locals. This means that AI translation systems can allow users to converse naturally with locals without having to study the language.

[0029] The AI ​​translation system according to this embodiment comprises a reception unit, a learning unit, an analysis unit, and a provision unit. The reception unit receives user input. User input includes, but is not limited to, text input, voice input, and image input. The reception unit may include, for example, a keyboard or touchscreen for receiving text input. The reception unit may also include a microphone or voice recognition technology for receiving voice input. Furthermore, the reception unit may also include a camera or image recognition technology for receiving image input. The learning unit learns based on data from social media and television. The learning unit collects, for example, social media posts and television subtitle data, and learns based on this data. The learning unit learns the latest words and expressions using machine learning algorithms and deep learning algorithms. For example, the learning unit collects social media posts and analyzes the meaning and context of words using natural language processing technology. The learning unit also collects television subtitle data and learns word usage and expressions based on this data. The analysis unit analyzes the input words. The analysis unit analyzes the grammar and meaning of the input words, for example, using natural language processing technology. The analysis unit performs grammatical and semantic analysis to understand the precise meaning of the input words. For example, the analysis unit analyzes the grammatical structure of the input words and identifies the subject, predicate, object, etc. of the sentence. The analysis unit also performs semantic analysis to understand the meaning of the input words. The provision unit provides the translation results. The provision unit provides the translation results, for example, in text format or audio format. The provision unit may be equipped with a display or speaker to show the translation results to the user. For example, the provision unit displays the translation results in text format on the display. The provision unit can also output the translation results in audio format from the speaker. Thus, the AI ​​translation system according to this embodiment can achieve natural translation by accepting user input, learning based on data from SNS and television, analyzing the input words, and providing translation results.

[0030] The reception unit receives user input. User input includes, but is not limited to, text input, voice input, and image input. The reception unit may, for example, be equipped with a keyboard or touchscreen for receiving text input, allowing users to easily input text. The reception unit may also be equipped with a microphone and speech recognition technology for receiving voice input. Speech recognition technology can convert user speech into text with high accuracy, which is particularly useful when the user's hands are occupied or when visual input is difficult. Furthermore, the reception unit may be equipped with a camera and image recognition technology for receiving image input. Image recognition technology can, for example, recognize handwritten characters or printed text and convert them into digital text. This allows users to input in various formats, improving the system's flexibility. The reception unit can integrate these input devices and centrally manage user input. For example, if a user provides voice input, the voice data is immediately converted into text and sent to the next processing step. Similarly, if an image is provided, it is converted into text data by image recognition technology and processed in the same way. This allows the reception desk to accommodate diverse user input formats and enables smooth data processing.

[0031] The learning unit learns based on data from social media and television. For example, it collects social media posts and television subtitle data, and uses this data to learn. Specifically, to collect social media posts, it uses APIs to obtain data in real time and analyzes the meaning and context of words using natural language processing technology. This allows it to quickly catch up on the latest trends and popular words. In addition, to collect television subtitle data, it regularly obtains data provided by broadcasting stations and learns about word usage and expressions based on this data. The learning unit uses machine learning algorithms and deep learning algorithms to analyze this data and update its language model. For example, it uses deep learning algorithms to train a model to generate appropriate translations according to context. This allows the learning unit to constantly learn the latest words and expressions and improve the accuracy of translations. Furthermore, the learning unit can collect and learn data to achieve natural translations that take context and cultural background into account, even when translating between different languages. For example, it can learn idioms and slang in a specific language and build a model to translate them appropriately. This allows the learning unit to provide users with more natural and accurate translations.

[0032] The analysis unit analyzes the input words. For example, it uses natural language processing techniques to analyze the grammar and meaning of the input words. Specifically, it performs grammatical and semantic analysis to understand the precise meaning of the input words. Grammatical analysis analyzes the grammatical structure of the input text, identifying the subject, predicate, object, etc. This allows for an accurate understanding of the sentence structure and lays the foundation for appropriate translation. Semantic analysis understands the meaning of the input words and generates appropriate translations based on the context. For example, the same word may have different meanings depending on the context; semantic analysis allows for the provision of contextually appropriate translations. Based on these analysis results, the analysis unit understands the precise meaning of the input words and sends them to the next processing step. Furthermore, the analysis unit can also analyze the sentiment and intent of the input words. For example, sentiment analysis can determine whether the input text conveys positive or negative emotions. Intent analysis can understand what the user wants to convey and provide appropriate translations accordingly. This allows the analysis unit to perform advanced analysis that considers not only the grammar and meaning of the input words, but also emotions and intentions, resulting in more natural and accurate translations.

[0033] The service provider delivers translation results. The service provider delivers translation results in various formats, such as text or audio. Specifically, it can be equipped with a display and speakers to show the translation results to the user. For example, the service provider can display the translation results in text format on the display, allowing the user to visually confirm the translation. The service provider can also output the translation results in audio format through the speakers, allowing the user to auditorily confirm the translation, which is particularly useful when visual confirmation is difficult. Furthermore, the service provider can send the translation results to other devices and applications. For example, the translation results can be shared with other users by sending them via email or messaging apps. The service provider can also be equipped with storage functionality to save translation results, allowing users to easily refer to past translation results. The service provider can integrate these functions to provide translation results to users in various formats. For example, if a user wants to check the translation results in text format, they can display it on the display; if they want to check it in audio format, they can output it through the speakers. They can also share the translation results with other users via email or messaging apps. This enables the service provider to provide flexible translation results tailored to user needs.

[0034] The AI ​​translation system includes an update unit that updates the training data. The update unit periodically updates the training data. For example, the update unit updates the training data based on a regular schedule. The update unit can also update the training data in real time. For example, the update unit updates the training data based on a regular schedule such as daily, weekly, or monthly. The update unit can also update the training data in real time whenever data from social media or television is updated. This allows the system to keep up with the latest words and expressions by regularly updating the training data. Some or all of the above processing in the update unit may be performed using AI, for example, or without AI. For example, the update unit can update the training data using an AI model that collects data from social media and television and updates the training data based on this data.

[0035] The AI ​​translation system includes a voice reception unit that accepts voice input. The voice reception unit accepts voice input. For example, the voice reception unit inputs the user's voice using a microphone. The voice reception unit can also convert the voice to text using speech recognition technology. For example, the voice reception unit inputs the voice the user speaks into the microphone and converts the voice to text using speech recognition technology. This allows the user to input the words they want to translate by voice by accepting voice input. Some or all of the above processing in the voice reception unit may be performed using AI, for example, or without AI. For example, the voice reception unit can input the voice data acquired by the microphone into a generating AI and have the generating AI perform the conversion from voice data to text data.

[0036] The reception desk can analyze the user's past input history and suggest the optimal input method. For example, the reception desk can automatically display words that the user has frequently entered in the past as suggestions. The reception desk can also prioritize suggesting input methods (voice, text, etc.) that the user has used in the past. Furthermore, the reception desk can predict and suggest words that the user will use at specific times of the day based on their past input history. In this way, the reception desk can suggest the optimal input method by analyzing the user's past input history. Some or all of the above processing in the reception desk may be performed using AI, for example, or not using AI. For example, the reception desk can input the user's past input history data into a generating AI and have the generating AI suggest the optimal input method.

[0037] The reception desk can customize the input method based on the user's current situation and environment when receiving input. For example, if the user is in a public place, the reception desk can suggest an input method that can be used in a quiet environment. The reception desk can also prioritize voice input if the user is on the move, allowing for quick input of the words to be translated. Furthermore, if the user is at home, the reception desk can provide detailed input options and suggest a customizable input method. This allows for the provision of a more appropriate input method by customizing it according to the user's current situation and environment. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's current situation and environment data into a generating AI and have the generating AI perform the customization of the input method.

[0038] The reception desk can prioritize inputs that are highly relevant, taking into account the user's geographical location. For example, if the user is in a specific region, the reception desk can prioritize inputs using words commonly used in that region. Similarly, if the user is traveling, the reception desk can prioritize inputs using words commonly used in tourist destinations. Furthermore, if the user is on a business trip, the reception desk can prioritize inputs using business terminology. This allows the reception desk to prioritize inputs that are highly relevant by considering the user's geographical location. Some or all of the above processing in the reception desk may be performed using AI, or not. For example, the reception desk can input the user's geographical location data into a generating AI and have the generating AI prioritize the most relevant inputs.

[0039] The reception unit can analyze the user's social media activity when receiving input and prioritize the acceptance of relevant input. For example, the reception unit can prioritize input of words that the user frequently uses on social media. Furthermore, if the user is posting about a specific topic, the reception unit can prioritize input of words related to that topic. Additionally, if the user is using a specific hashtag, the reception unit can prioritize input of words related to that hashtag. This allows the reception unit to prioritize the acceptance of relevant input by analyzing the user's social media activity. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input the user's social media activity data into a generating AI and have the generating AI prioritize relevant inputs.

[0040] The learning unit can optimize the learning algorithm by referring to past learning data during the learning process. For example, the learning unit can select the optimal learning algorithm based on past learning data. The learning unit can also analyze past learning data and adjust the parameters of the learning algorithm. Furthermore, the learning unit can improve the accuracy of the learning algorithm by referring to past learning data. In this way, the learning algorithm can be optimized by referring to past learning data. Some or all of the above processes in the learning unit may be performed using AI, for example, or without using AI. For example, the learning unit can input past learning data into a generating AI and have the generating AI perform the optimization of the learning algorithm.

[0041] The learning unit can improve its learning accuracy by using datasets specific to particular languages ​​or cultures during training. For example, the learning unit can learn using a dataset specific to the expressions used by Japanese business people. It can also learn using a dataset specific to youth slang. Furthermore, the learning unit can learn using datasets specific to particular regions or cultures. This allows for improved learning accuracy by using datasets specific to particular languages ​​or cultures. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input a dataset specific to a particular language or culture into a generating AI and have the generating AI perform the improvement of learning accuracy.

[0042] The learning unit can learn by adding data from blogs and news sites in addition to data from social media and television during the learning process. For example, the learning unit can learn the latest words and expressions based on social media data. It can also learn common words and expressions based on television data. Furthermore, it can learn specialized words and expressions based on data from blogs and news sites. This allows the learning unit to learn a wider variety of words and expressions by learning from blogs and news sites in addition to social media and television data. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input data from blogs and news sites into a generating AI and have the generating AI perform the addition of learning content.

[0043] The learning unit can update the learning content in real time by incorporating user feedback during the learning process. For example, the learning unit updates the learning content in real time based on feedback provided by the user. The learning unit can also analyze user feedback and adjust the learning algorithm. Furthermore, the learning unit can add or delete learning data by referring to user feedback. This allows for rapid updates of the learning content by incorporating user feedback in real time. Some or all of the above processes in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input user feedback data into a generating AI and have the generating AI perform the update of the learning content.

[0044] The analysis unit can improve analysis accuracy by considering the context of the input words during analysis. For example, the analysis unit can analyze by considering the context before and after the input words. The analysis unit can also analyze by considering the category of the input words. Furthermore, the analysis unit can analyze by considering the frequency of use of the input words. In this way, the analysis accuracy can be improved by considering the context of the input words. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input contextual data of the input words into a generating AI and have the generating AI perform the improvement of analysis accuracy.

[0045] The analysis unit can apply different analysis methods depending on the category of the input words during analysis. For example, the analysis unit can apply an analysis method specialized for business terminology to business expressions. It can also apply an analysis method specialized for youth slang to youth slang. Furthermore, it can apply a general analysis method to general words. By applying different analysis methods depending on the category of the input words, the accuracy of the analysis can be improved. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the category data of the input words into a generating AI and have the generating AI perform the application of the analysis method.

[0046] The analysis unit can perform analysis while considering the geographical background of the input words. For example, if the input words are used in a specific region, the analysis unit will consider the background of that region. Furthermore, if the input words are related to a specific culture, the analysis unit can also consider the background of that culture. In addition, if the input words are related to a specific event, the analysis unit can also consider the background of that event. This allows for improved analysis accuracy by considering the geographical background of the input words. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input geographical background data of the input words into a generating AI and have the generating AI perform the task of improving analysis accuracy.

[0047] The analysis unit can improve the accuracy of its analysis by referring to relevant literature for the input words during the analysis process. For example, the analysis unit can analyze by referring to literature related to the input words. It can also analyze by referring to research papers related to the input words. Furthermore, the analysis unit can analyze by referring to news articles related to the input words. In this way, the accuracy of the analysis can be improved by referring to literature related to the input words. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input data on relevant literature for the input words into a generating AI and have the generating AI perform the improvement of the analysis accuracy.

[0048] The service provider can adjust the level of detail displayed based on the importance of the translation result at the time of delivery. For example, in the case of an important business document, the service provider can provide a detailed translation result. In the case of everyday conversation, the service provider can also provide a concise translation result. Furthermore, in the case of a speech in a formal setting, the service provider can provide an accurate and detailed translation result. This allows important information to be displayed preferentially by adjusting the level of detail displayed based on the importance of the translation result. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input the importance data of the translation result into a generating AI and have the generating AI perform the adjustment of the level of detail displayed.

[0049] The service provider can apply different display algorithms depending on the category of the translation result at the time of delivery. For example, in the case of business documents, the service provider can apply a display algorithm specialized for business terminology. It can also apply a display algorithm specialized for youth slang. Furthermore, it can apply a general display algorithm for general everyday conversation. This allows for a more appropriate display method by applying different display algorithms depending on the category of the translation result. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can input the category data of the translation result into a generating AI and have the generating AI perform the application of the display algorithm.

[0050] The service provider can determine the display priority based on the submission timing of the translation results at the time of delivery. For example, the service provider will display urgent translation requests with the highest priority. The service provider can also display regular translation requests with normal priority. Furthermore, the service provider can display long-term translation requests later. This allows for the priority display of urgent translation requests by determining the display priority based on the submission timing of the translation results. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input translation result submission timing data into a generating AI and have the generating AI perform the determination of the display priority.

[0051] The delivery unit can adjust the display order based on the relevance of the translation results at the time of delivery. For example, the delivery unit can display important business documents first. It can also display everyday conversations later. Furthermore, it can prioritize displaying speeches in formal settings. This allows important information to be displayed preferentially by adjusting the display order based on the relevance of the translation results. Some or all of the above processing in the delivery unit may be performed using AI, for example, or not using AI. For example, the delivery unit can input relevance data of the translation results into a generating AI and have the generating AI perform the adjustment of the display order.

[0052] The update unit can select the optimal update method by referring to past update history during an update. For example, the update unit selects the optimal update method based on past update history. The update unit can also analyze past update history and adjust the parameters of the update method. Furthermore, the update unit can improve the accuracy of the update method by referring to past update history. This allows the optimal update method to be selected by referring to past update history. Some or all of the above processing in the update unit may be performed using AI, for example, or without AI. For example, the update unit can input past update history data into a generating AI and have the generating AI perform the selection of the update method.

[0053] The update unit can update by adding data from blogs and news sites in addition to data from social media and television. For example, the update unit can update based on social media data to add the latest words and expressions. It can also update based on television data to add common words and expressions. Furthermore, it can update based on blog and news site data to add specialized words and expressions. In this way, by adding data from blogs and news sites in addition to social media and television data, a wider variety of words and expressions can be updated. Some or all of the above processing in the update unit may be performed using AI, for example, or not using AI. For example, the update unit can input data from blogs and news sites into a generating AI and have the generating AI execute the addition of update content.

[0054] The voice reception unit can suggest the optimal input method by referring to the user's past voice input history when voice input is performed. For example, the voice reception unit may prioritize suggesting voice input methods that the user has frequently used in the past. The voice reception unit can also predict and suggest words that the user will use at specific times of day based on the user's past voice input history. Furthermore, the voice reception unit can analyze the user's past voice input history and suggest the optimal voice input method. In this way, the optimal voice input method can be suggested by referring to the user's past voice input history. Some or all of the above processing in the voice reception unit may be performed using AI, for example, or without AI. For example, the voice reception unit can input the user's past voice input history data into a generating AI and have the generating AI suggest the optimal voice input method.

[0055] The voice reception unit can prioritize receiving voice input that is highly relevant, taking into account the user's geographical location. For example, if the user is in a specific region, the voice reception unit can prioritize voice input of words commonly used in that region. Furthermore, if the user is traveling, the voice reception unit can prioritize voice input of words commonly used in tourist destinations. Additionally, if the user is on a business trip, the voice reception unit can prioritize voice input of business terminology. This allows the system to prioritize receiving voice input that is highly relevant by considering the user's geographical location. Some or all of the above processing in the voice reception unit may be performed using AI, for example, or without AI. For example, the voice reception unit can input the user's geographical location data into a generating AI and have the generating AI prioritize highly relevant voice inputs.

[0056] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0057] AI translation systems can provide relevant cultural background information based on user input. For example, if a user inputs "Let's all do our best in baseball," the system can provide information about the Japanese business culture and the importance of teamwork behind this phrase. Similarly, if a user inputs "Haoi," the system can provide information about how this word is used among young people and its nuances. Furthermore, if a user inputs a specific regional language, the system can provide information about the culture and history of that region. This allows users not only to obtain translation results but also to deepen their understanding of the cultural context behind the words.

[0058] The AI ​​translation system can provide relevant images and videos based on user input. For example, if a user inputs "Let's all do our best together," the system can provide images and videos depicting Japanese business scenes and teamwork. Similarly, if a user inputs "Haoi," the system can provide images and videos showing young people using this word. Furthermore, if a user inputs a specific regional language, the system can provide images and videos showcasing the scenery and culture of that region. This makes the translation results easier for users to understand visually, leading to a deeper comprehension.

[0059] The AI ​​translation system can provide relevant music and audio based on the user's input. For example, if a user inputs "Let's all do our best together," the system can provide cheering songs and audio messages commonly used in Japanese business settings. If a user inputs "Haoi," the system can provide audio and music depicting young people using this word. Furthermore, if a user inputs a specific regional language, the system can introduce traditional music and audio from that region. This makes the translation results easier for the user to understand aurally, leading to a deeper comprehension.

[0060] The AI ​​translation system can provide relevant news articles and blog posts based on user input. For example, if a user inputs "Let's all do our best together," the system can provide the latest news articles and blog posts on Japanese business culture and teamwork. If the user inputs "Haoi," the system can provide articles and blog posts on youth slang. Furthermore, if the user inputs a specific regional dialect, the system can provide the latest news and blog posts related to that region. This allows users to obtain up-to-date information through translation results and gain a deeper understanding.

[0061] The AI ​​translation system can recommend relevant books and articles based on user input. For example, if a user inputs "Let's all work together," the system can recommend books and articles on Japanese business culture and teamwork. If the user inputs "Haoi," the system can recommend research and books on youth slang. Furthermore, if the user inputs a specific regional dialect, the system can recommend books and articles related to that region. This allows users to gain deeper knowledge through the translation results and deepen their academic understanding.

[0062] The following briefly describes the processing flow for example form 1.

[0063] Step 1: The reception area receives user input. User input includes text input, voice input, and image input. The reception area is equipped with a keyboard or touchscreen for receiving text input, a microphone and voice recognition technology for receiving voice input, and a camera and image recognition technology for receiving image input. Step 2: The learning unit learns based on data from social media and television. The learning unit collects social media posts and television subtitle data, and uses machine learning algorithms and deep learning algorithms to learn the latest words and expressions. For example, it uses natural language processing techniques to analyze the meaning and context of words. Step 3: The analysis unit analyzes the input words. The analysis unit uses natural language processing technology to analyze the grammar and meaning of the input words, performing grammatical and semantic analysis. For example, it identifies the subject, predicate, and object of a sentence and understands the meaning of the input words. Step 4: The service provider provides the translation results. The service provider provides the translation results in text or audio format and is equipped with a display and speakers to show them to the user. For example, the translation results may be displayed on the display in text format and output from the speakers in audio format.

[0064] (Example of form 2) The AI ​​translation system according to an embodiment of the present invention is a system that provides a translation service that allows users to converse naturally with locals without having to learn the language. This AI translation system works by having the user input the words they want to translate, and the AI ​​performs a natural translation based on learning data of words used on social media and television. This learning data is regularly updated, allowing the system to handle the latest words and expressions. For example, in Japan, unique business phrases such as "Let's all work together" and slang terms like "Haoi" are constantly emerging. The AI ​​learns these words and expressions and translates them in a way that foreigners can naturally understand. For example, the user inputs a business phrase like "Let's all work together" or slang terms like "Haoi." This information is input to the AI. Next, the AI ​​analyzes the input words and performs a natural translation based on learning data of words used on social media and television. By regularly updating this learning data, the AI ​​can handle the latest words and expressions. For example, it translates the phrase "Let's all work together" into an expression that is easy for foreigners to understand. This mechanism allows users to converse naturally with locals without having to learn the language. For example, AI can naturally translate words and expressions that are difficult for foreigners to fully understand, even after extensive Japanese study, enabling them to communicate smoothly with locals. This means that AI translation systems can allow users to converse naturally with locals without having to study the language.

[0065] The AI ​​translation system according to this embodiment comprises a reception unit, a learning unit, an analysis unit, and a provision unit. The reception unit receives user input. User input includes, but is not limited to, text input, voice input, and image input. The reception unit may include, for example, a keyboard or touchscreen for receiving text input. The reception unit may also include a microphone or voice recognition technology for receiving voice input. Furthermore, the reception unit may also include a camera or image recognition technology for receiving image input. The learning unit learns based on data from social media and television. The learning unit collects, for example, social media posts and television subtitle data, and learns based on this data. The learning unit learns the latest words and expressions using machine learning algorithms and deep learning algorithms. For example, the learning unit collects social media posts and analyzes the meaning and context of words using natural language processing technology. The learning unit also collects television subtitle data and learns word usage and expressions based on this data. The analysis unit analyzes the input words. The analysis unit analyzes the grammar and meaning of the input words, for example, using natural language processing technology. The analysis unit performs grammatical and semantic analysis to understand the precise meaning of the input words. For example, the analysis unit analyzes the grammatical structure of the input words and identifies the subject, predicate, object, etc. of the sentence. The analysis unit also performs semantic analysis to understand the meaning of the input words. The provision unit provides the translation results. The provision unit provides the translation results, for example, in text format or audio format. The provision unit may be equipped with a display or speaker to show the translation results to the user. For example, the provision unit displays the translation results in text format on the display. The provision unit can also output the translation results in audio format from the speaker. Thus, the AI ​​translation system according to this embodiment can achieve natural translation by accepting user input, learning based on data from SNS and television, analyzing the input words, and providing translation results.

[0066] The reception unit receives user input. User input includes, but is not limited to, text input, voice input, and image input. The reception unit may, for example, be equipped with a keyboard or touchscreen for receiving text input, allowing users to easily input text. The reception unit may also be equipped with a microphone and speech recognition technology for receiving voice input. Speech recognition technology can convert user speech into text with high accuracy, which is particularly useful when the user's hands are occupied or when visual input is difficult. Furthermore, the reception unit may be equipped with a camera and image recognition technology for receiving image input. Image recognition technology can, for example, recognize handwritten characters or printed text and convert them into digital text. This allows users to input in various formats, improving the system's flexibility. The reception unit can integrate these input devices and centrally manage user input. For example, if a user provides voice input, the voice data is immediately converted into text and sent to the next processing step. Similarly, if an image is provided, it is converted into text data by image recognition technology and processed in the same way. This allows the reception desk to accommodate diverse user input formats and enables smooth data processing.

[0067] The learning unit learns based on data from social media and television. For example, it collects social media posts and television subtitle data, and uses this data to learn. Specifically, to collect social media posts, it uses APIs to obtain data in real time and analyzes the meaning and context of words using natural language processing technology. This allows it to quickly catch up on the latest trends and popular words. In addition, to collect television subtitle data, it regularly obtains data provided by broadcasting stations and learns about word usage and expressions based on this data. The learning unit uses machine learning algorithms and deep learning algorithms to analyze this data and update its language model. For example, it uses deep learning algorithms to train a model to generate appropriate translations according to context. This allows the learning unit to constantly learn the latest words and expressions and improve the accuracy of translations. Furthermore, the learning unit can collect and learn data to achieve natural translations that take context and cultural background into account, even when translating between different languages. For example, it can learn idioms and slang in a specific language and build a model to translate them appropriately. This allows the learning unit to provide users with more natural and accurate translations.

[0068] The analysis unit analyzes the input words. For example, it uses natural language processing techniques to analyze the grammar and meaning of the input words. Specifically, it performs grammatical and semantic analysis to understand the precise meaning of the input words. Grammatical analysis analyzes the grammatical structure of the input text, identifying the subject, predicate, object, etc. This allows for an accurate understanding of the sentence structure and lays the foundation for appropriate translation. Semantic analysis understands the meaning of the input words and generates appropriate translations based on the context. For example, the same word may have different meanings depending on the context; semantic analysis allows for the provision of contextually appropriate translations. Based on these analysis results, the analysis unit understands the precise meaning of the input words and sends them to the next processing step. Furthermore, the analysis unit can also analyze the sentiment and intent of the input words. For example, sentiment analysis can determine whether the input text conveys positive or negative emotions. Intent analysis can understand what the user wants to convey and provide appropriate translations accordingly. This allows the analysis unit to perform advanced analysis that considers not only the grammar and meaning of the input words, but also emotions and intentions, resulting in more natural and accurate translations.

[0069] The service provider delivers translation results. The service provider delivers translation results in various formats, such as text or audio. Specifically, it can be equipped with a display and speakers to show the translation results to the user. For example, the service provider can display the translation results in text format on the display, allowing the user to visually confirm the translation. The service provider can also output the translation results in audio format through the speakers, allowing the user to auditorily confirm the translation, which is particularly useful when visual confirmation is difficult. Furthermore, the service provider can send the translation results to other devices and applications. For example, the translation results can be shared with other users by sending them via email or messaging apps. The service provider can also be equipped with storage functionality to save translation results, allowing users to easily refer to past translation results. The service provider can integrate these functions to provide translation results to users in various formats. For example, if a user wants to check the translation results in text format, they can display it on the display; if they want to check it in audio format, they can output it through the speakers. They can also share the translation results with other users via email or messaging apps. This enables the service provider to provide flexible translation results tailored to user needs.

[0070] The AI ​​translation system includes an update unit that updates the training data. The update unit periodically updates the training data. For example, the update unit updates the training data based on a regular schedule. The update unit can also update the training data in real time. For example, the update unit updates the training data based on a regular schedule such as daily, weekly, or monthly. The update unit can also update the training data in real time whenever data from social media or television is updated. This allows the system to keep up with the latest words and expressions by regularly updating the training data. Some or all of the above processing in the update unit may be performed using AI, for example, or without AI. For example, the update unit can update the training data using an AI model that collects data from social media and television and updates the training data based on this data.

[0071] The AI ​​translation system includes a voice reception unit that accepts voice input. The voice reception unit accepts voice input. For example, the voice reception unit inputs the user's voice using a microphone. The voice reception unit can also convert the voice to text using speech recognition technology. For example, the voice reception unit inputs the voice the user speaks into the microphone and converts the voice to text using speech recognition technology. This allows the user to input the words they want to translate by voice by accepting voice input. Some or all of the above processing in the voice reception unit may be performed using AI, for example, or without AI. For example, the voice reception unit can input the voice data acquired by the microphone into a generating AI and have the generating AI perform the conversion from voice data to text data.

[0072] The reception unit can estimate the user's emotions and adjust the input interface based on the estimated emotions. For example, if the user is stressed, the reception unit can provide a simple interface and minimize the input steps. If the user is relaxed, the reception unit can also provide detailed input options and suggest customizable input methods. Furthermore, if the user is in a hurry, the reception unit can prioritize voice input, allowing for quick input of the words to be translated. This allows for a user-friendly interface by adjusting the input interface according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception unit may be performed using AI or not. For example, the reception unit can input user image data captured by a camera into a generative AI and have the generative AI perform the user's emotion estimation.

[0073] The reception desk can analyze the user's past input history and suggest the optimal input method. For example, the reception desk can automatically display words that the user has frequently entered in the past as suggestions. The reception desk can also prioritize suggesting input methods (voice, text, etc.) that the user has used in the past. Furthermore, the reception desk can predict and suggest words that the user will use at specific times of the day based on their past input history. In this way, the reception desk can suggest the optimal input method by analyzing the user's past input history. Some or all of the above processing in the reception desk may be performed using AI, for example, or not using AI. For example, the reception desk can input the user's past input history data into a generating AI and have the generating AI suggest the optimal input method.

[0074] The reception desk can customize the input method based on the user's current situation and environment when receiving input. For example, if the user is in a public place, the reception desk can suggest an input method that can be used in a quiet environment. The reception desk can also prioritize voice input if the user is on the move, allowing for quick input of the words to be translated. Furthermore, if the user is at home, the reception desk can provide detailed input options and suggest a customizable input method. This allows for the provision of a more appropriate input method by customizing it according to the user's current situation and environment. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's current situation and environment data into a generating AI and have the generating AI perform the customization of the input method.

[0075] The reception desk can estimate the user's emotions and prioritize input based on those emotions. For example, if the user is nervous, the reception desk may prioritize inputting important words. If the user is relaxed, the reception desk may also provide detailed input options and suggest customizable input methods. Furthermore, if the user is in a hurry, the reception desk may prioritize voice input, allowing them to quickly input the words they want translated. This ensures that important information is prioritized by prioritizing input according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception desk may be performed using AI or not. For example, the reception desk may input user image data captured by a camera into a generative AI and have the generative AI perform emotion estimation.

[0076] The reception desk can prioritize inputs that are highly relevant, taking into account the user's geographical location. For example, if the user is in a specific region, the reception desk can prioritize inputs using words commonly used in that region. Similarly, if the user is traveling, the reception desk can prioritize inputs using words commonly used in tourist destinations. Furthermore, if the user is on a business trip, the reception desk can prioritize inputs using business terminology. This allows the reception desk to prioritize inputs that are highly relevant by considering the user's geographical location. Some or all of the above processing in the reception desk may be performed using AI, or not. For example, the reception desk can input the user's geographical location data into a generating AI and have the generating AI prioritize the most relevant inputs.

[0077] The reception unit can analyze the user's social media activity when receiving input and prioritize the acceptance of relevant input. For example, the reception unit can prioritize input of words that the user frequently uses on social media. Furthermore, if the user is posting about a specific topic, the reception unit can prioritize input of words related to that topic. Additionally, if the user is using a specific hashtag, the reception unit can prioritize input of words related to that hashtag. This allows the reception unit to prioritize the acceptance of relevant input by analyzing the user's social media activity. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input the user's social media activity data into a generating AI and have the generating AI prioritize relevant inputs.

[0078] The learning unit can estimate the user's emotions and select training data based on the estimated user emotions. For example, if the user is relaxed, the learning unit will select words used in a relaxed state as training data. Similarly, if the user is tense, the learning unit can select words used in a tense state. Furthermore, if the user is excited, the learning unit can select words used in an excited state as training data. This allows for the learning of more appropriate data by selecting training data according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI may be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processes in the learning unit may be performed using AI, or not. For example, the learning unit can input user emotion data into a generative AI and have the generative AI perform the selection of training data.

[0079] The learning unit can optimize the learning algorithm by referring to past learning data during the learning process. For example, the learning unit can select the optimal learning algorithm based on past learning data. The learning unit can also analyze past learning data and adjust the parameters of the learning algorithm. Furthermore, the learning unit can improve the accuracy of the learning algorithm by referring to past learning data. In this way, the learning algorithm can be optimized by referring to past learning data. Some or all of the above processes in the learning unit may be performed using AI, for example, or without using AI. For example, the learning unit can input past learning data into a generating AI and have the generating AI perform the optimization of the learning algorithm.

[0080] The learning unit can improve its learning accuracy by using datasets specific to particular languages ​​or cultures during training. For example, the learning unit can learn using a dataset specific to the expressions used by Japanese business people. It can also learn using a dataset specific to youth slang. Furthermore, the learning unit can learn using datasets specific to particular regions or cultures. This allows for improved learning accuracy by using datasets specific to particular languages ​​or cultures. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input a dataset specific to a particular language or culture into a generating AI and have the generating AI perform the improvement of learning accuracy.

[0081] The learning unit can estimate the user's emotions and adjust the learning frequency based on the estimated emotions. For example, if the user is relaxed, the learning unit can set a low learning frequency. If the user is stressed, the learning unit can also set a high learning frequency. Furthermore, if the user is excited, the learning unit can set a medium learning frequency. By adjusting the learning frequency according to the user's emotions, more effective learning becomes possible. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input user emotion data into the generative AI and have the generative AI adjust the learning frequency.

[0082] The learning unit can learn by adding data from blogs and news sites in addition to data from social media and television during the learning process. For example, the learning unit can learn the latest words and expressions based on social media data. It can also learn common words and expressions based on television data. Furthermore, it can learn specialized words and expressions based on data from blogs and news sites. This allows the learning unit to learn a wider variety of words and expressions by learning from blogs and news sites in addition to social media and television data. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input data from blogs and news sites into a generating AI and have the generating AI perform the addition of learning content.

[0083] The learning unit can update the learning content in real time by incorporating user feedback during the learning process. For example, the learning unit updates the learning content in real time based on feedback provided by the user. The learning unit can also analyze user feedback and adjust the learning algorithm. Furthermore, the learning unit can add or delete learning data by referring to user feedback. This allows for rapid updates of the learning content by incorporating user feedback in real time. Some or all of the above processes in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can input user feedback data into a generating AI and have the generating AI perform the update of the learning content.

[0084] The analysis unit can estimate the user's emotions and adjust the analysis algorithm based on the estimated emotions. For example, if the user is relaxed, the analysis unit may use an algorithm that performs a detailed analysis. If the user is tense, the analysis unit may also use an algorithm that performs a concise analysis. Furthermore, if the user is excited, the analysis unit may use an algorithm that provides visually stimulating analysis results. By adjusting the analysis algorithm according to the user's emotions, more appropriate analysis results can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input user emotion data into the generative AI and have the generative AI adjust the analysis algorithm.

[0085] The analysis unit can improve analysis accuracy by considering the context of the input words during analysis. For example, the analysis unit can analyze by considering the context before and after the input words. The analysis unit can also analyze by considering the category of the input words. Furthermore, the analysis unit can analyze by considering the frequency of use of the input words. In this way, the analysis accuracy can be improved by considering the context of the input words. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input contextual data of the input words into a generating AI and have the generating AI perform the improvement of analysis accuracy.

[0086] The analysis unit can apply different analysis methods depending on the category of the input words during analysis. For example, the analysis unit can apply an analysis method specialized for business terminology to business expressions. It can also apply an analysis method specialized for youth slang to youth slang. Furthermore, it can apply a general analysis method to general words. By applying different analysis methods depending on the category of the input words, the accuracy of the analysis can be improved. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the category data of the input words into a generating AI and have the generating AI perform the application of the analysis method.

[0087] The analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated emotions. For example, if the user is nervous, the analysis unit can provide a simple and highly visible display method. If the user is relaxed, the analysis unit can also provide a display method that includes detailed information. Furthermore, if the user is in a hurry, the analysis unit can provide a concise display method. In this way, by adjusting the display method of the analysis results according to the user's emotions, a more appropriate display method can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's emotion data into the generative AI and have the generative AI perform the adjustment of the display method.

[0088] The analysis unit can perform analysis while considering the geographical background of the input words. For example, if the input words are used in a specific region, the analysis unit will consider the background of that region. Furthermore, if the input words are related to a specific culture, the analysis unit can also consider the background of that culture. In addition, if the input words are related to a specific event, the analysis unit can also consider the background of that event. This allows for improved analysis accuracy by considering the geographical background of the input words. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input geographical background data of the input words into a generating AI and have the generating AI perform the task of improving analysis accuracy.

[0089] The analysis unit can improve the accuracy of its analysis by referring to relevant literature for the input words during the analysis process. For example, the analysis unit can analyze by referring to literature related to the input words. It can also analyze by referring to research papers related to the input words. Furthermore, the analysis unit can analyze by referring to news articles related to the input words. In this way, the accuracy of the analysis can be improved by referring to literature related to the input words. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input data on relevant literature for the input words into a generating AI and have the generating AI perform the improvement of the analysis accuracy.

[0090] The service provider can estimate the user's emotions and adjust the way the translation results are presented based on the estimated emotions. For example, if the user is relaxed, the service provider can provide a detailed translation. If the user is tense, the service provider can also provide a concise translation. Furthermore, if the user is excited, the service provider can provide a visually stimulating translation. By adjusting the way the translation results are presented according to the user's emotions, more appropriate translation results can be provided. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can input user emotion data into a generative AI and have the generative AI adjust the way the translation results are presented.

[0091] The service provider can adjust the level of detail displayed based on the importance of the translation result at the time of delivery. For example, in the case of an important business document, the service provider can provide a detailed translation result. In the case of everyday conversation, the service provider can also provide a concise translation result. Furthermore, in the case of a speech in a formal setting, the service provider can provide an accurate and detailed translation result. This allows important information to be displayed preferentially by adjusting the level of detail displayed based on the importance of the translation result. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input the importance data of the translation result into a generating AI and have the generating AI perform the adjustment of the level of detail displayed.

[0092] The service provider can apply different display algorithms depending on the category of the translation result at the time of delivery. For example, in the case of business documents, the service provider can apply a display algorithm specialized for business terminology. It can also apply a display algorithm specialized for youth slang. Furthermore, it can apply a general display algorithm for general everyday conversation. This allows for a more appropriate display method by applying different display algorithms depending on the category of the translation result. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can input the category data of the translation result into a generating AI and have the generating AI perform the application of the display algorithm.

[0093] The service provider can estimate the user's emotions and adjust the length of the translation result based on the estimated emotions. For example, if the user is relaxed, the service provider can provide a detailed translation result. If the user is stressed, the service provider can also provide a concise translation result. Furthermore, if the user is in a hurry, the service provider can provide a short, to-the-point translation result. By adjusting the length of the translation result according to the user's emotions, a more appropriate translation result can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input user emotion data into the generative AI and have the generative AI adjust the length of the translation result.

[0094] The service provider can determine the display priority based on the submission timing of the translation results at the time of delivery. For example, the service provider will display urgent translation requests with the highest priority. The service provider can also display regular translation requests with normal priority. Furthermore, the service provider can display long-term translation requests later. This allows for the priority display of urgent translation requests by determining the display priority based on the submission timing of the translation results. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input translation result submission timing data into a generating AI and have the generating AI perform the determination of the display priority.

[0095] The delivery unit can adjust the display order based on the relevance of the translation results at the time of delivery. For example, the delivery unit can display important business documents first. It can also display everyday conversations later. Furthermore, it can prioritize displaying speeches in formal settings. This allows important information to be displayed preferentially by adjusting the display order based on the relevance of the translation results. Some or all of the above processing in the delivery unit may be performed using AI, for example, or not using AI. For example, the delivery unit can input relevance data of the translation results into a generating AI and have the generating AI perform the adjustment of the display order.

[0096] The update unit can estimate the user's emotions and adjust the timing of updates based on the estimated emotions. For example, if the user is relaxed, the update unit can set a lower update frequency. If the user is stressed, the update unit can set a higher update frequency. Furthermore, if the user is excited, the update unit can set a moderate update frequency. By adjusting the timing of updates according to the user's emotions, updates can be performed at a more appropriate time. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the update unit may be performed using AI, or not using AI. For example, the update unit can input user emotion data into a generative AI and have the generative AI adjust the timing of updates.

[0097] The update unit can select the optimal update method by referring to past update history during an update. For example, the update unit selects the optimal update method based on past update history. The update unit can also analyze past update history and adjust the parameters of the update method. Furthermore, the update unit can improve the accuracy of the update method by referring to past update history. This allows the optimal update method to be selected by referring to past update history. Some or all of the above processing in the update unit may be performed using AI, for example, or without AI. For example, the update unit can input past update history data into a generating AI and have the generating AI perform the selection of the update method.

[0098] The update unit can estimate the user's emotions and determine the priority of updates based on the estimated emotions. For example, if the user is relaxed, the update unit may set a low priority for updates. Conversely, if the user is stressed, the update unit may set a high priority for updates. Furthermore, if the user is excited, the update unit may set a medium priority for updates. This allows important updates to be prioritized by determining the priority of updates according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the update unit may be performed using AI, or not using AI. For example, the update unit can input user emotion data into a generative AI and have the generative AI determine the priority of updates.

[0099] The update unit can update by adding data from blogs and news sites in addition to data from social media and television. For example, the update unit can update based on social media data to add the latest words and expressions. It can also update based on television data to add common words and expressions. Furthermore, it can update based on blog and news site data to add specialized words and expressions. In this way, by adding data from blogs and news sites in addition to social media and television data, a wider variety of words and expressions can be updated. Some or all of the above processing in the update unit may be performed using AI, for example, or not using AI. For example, the update unit can input data from blogs and news sites into a generating AI and have the generating AI execute the addition of update content.

[0100] The voice reception unit can estimate the user's emotions and adjust the voice input interface based on the estimated emotions. For example, if the user is relaxed, the voice reception unit can provide detailed voice input options. If the user is tense, it can also provide a simple voice input interface. Furthermore, if the user is in a hurry, it can provide an interface that allows for quick voice input. This allows for a more user-friendly voice input interface by adjusting it according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the voice reception unit may be performed using AI or not. For example, the voice reception unit can input user emotion data into a generative AI and have the generative AI adjust the voice input interface.

[0101] The voice reception unit can suggest the optimal input method by referring to the user's past voice input history when voice input is performed. For example, the voice reception unit may prioritize suggesting voice input methods that the user has frequently used in the past. The voice reception unit can also predict and suggest words that the user will use at specific times of day based on the user's past voice input history. Furthermore, the voice reception unit can analyze the user's past voice input history and suggest the optimal voice input method. In this way, the optimal voice input method can be suggested by referring to the user's past voice input history. Some or all of the above processing in the voice reception unit may be performed using AI, for example, or without AI. For example, the voice reception unit can input the user's past voice input history data into a generating AI and have the generating AI suggest the optimal voice input method.

[0102] The voice reception unit can estimate the user's emotions and prioritize voice input based on the estimated emotions. For example, if the user is relaxed, the voice reception unit will prioritize detailed voice input. It can also prioritize important voice input if the user is stressed. Furthermore, if the user is in a hurry, it can prioritize content that can be quickly voice-inputted. This allows for prioritizing important content based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the voice reception unit may be performed using AI or not. For example, the voice reception unit can input user emotion data into a generative AI and have the generative AI determine the priority of voice input.

[0103] The voice reception unit can prioritize receiving voice input that is highly relevant, taking into account the user's geographical location. For example, if the user is in a specific region, the voice reception unit can prioritize voice input of words commonly used in that region. Furthermore, if the user is traveling, the voice reception unit can prioritize voice input of words commonly used in tourist destinations. Additionally, if the user is on a business trip, the voice reception unit can prioritize voice input of business terminology. This allows the system to prioritize receiving voice input that is highly relevant by considering the user's geographical location. Some or all of the above processing in the voice reception unit may be performed using AI, for example, or without AI. For example, the voice reception unit can input the user's geographical location data into a generating AI and have the generating AI prioritize highly relevant voice inputs.

[0104] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0105] AI translation systems can provide relevant cultural background information based on user input. For example, if a user inputs "Let's all do our best in baseball," the system can provide information about the Japanese business culture and the importance of teamwork behind this phrase. Similarly, if a user inputs "Haoi," the system can provide information about how this word is used among young people and its nuances. Furthermore, if a user inputs a specific regional language, the system can provide information about the culture and history of that region. This allows users not only to obtain translation results but also to deepen their understanding of the cultural context behind the words.

[0106] The AI ​​translation system can provide relevant images and videos based on user input. For example, if a user inputs "Let's all do our best together," the system can provide images and videos depicting Japanese business scenes and teamwork. Similarly, if a user inputs "Haoi," the system can provide images and videos showing young people using this word. Furthermore, if a user inputs a specific regional language, the system can provide images and videos showcasing the scenery and culture of that region. This makes the translation results easier for users to understand visually, leading to a deeper comprehension.

[0107] The AI ​​translation system can provide relevant music and audio based on the user's input. For example, if a user inputs "Let's all do our best together," the system can provide cheering songs and audio messages commonly used in Japanese business settings. If a user inputs "Haoi," the system can provide audio and music depicting young people using this word. Furthermore, if a user inputs a specific regional language, the system can introduce traditional music and audio from that region. This makes the translation results easier for the user to understand aurally, leading to a deeper comprehension.

[0108] The AI ​​translation system can provide relevant news articles and blog posts based on user input. For example, if a user inputs "Let's all do our best together," the system can provide the latest news articles and blog posts on Japanese business culture and teamwork. If the user inputs "Haoi," the system can provide articles and blog posts on youth slang. Furthermore, if the user inputs a specific regional dialect, the system can provide the latest news and blog posts related to that region. This allows users to obtain up-to-date information through translation results and gain a deeper understanding.

[0109] The AI ​​translation system can recommend relevant books and articles based on user input. For example, if a user inputs "Let's all work together," the system can recommend books and articles on Japanese business culture and teamwork. If the user inputs "Haoi," the system can recommend research and books on youth slang. Furthermore, if the user inputs a specific regional dialect, the system can recommend books and articles related to that region. This allows users to gain deeper knowledge through the translation results and deepen their academic understanding.

[0110] AI translation systems can estimate a user's emotions and adjust the tone of the translation based on those emotions. For example, if a user is stressed, the system can provide a relaxing and gentle translation. If a user is excited, the system can provide a more energetic translation. Furthermore, if a user is sad, the system can provide a comforting translation. This allows for translations to be delivered in an appropriate tone according to the user's emotions, thereby improving user satisfaction.

[0111] AI translation systems can estimate a user's emotions and adjust the level of detail in the translation based on that estimation. For example, if the user is relaxed, the system can provide a detailed translation. If the user is stressed, it can provide a concise translation. Furthermore, if the user is in a hurry, it can provide a short, to-the-point translation. This allows the system to provide translations with the appropriate level of detail according to the user's emotions, thus providing a service that meets the user's needs.

[0112] AI translation systems can estimate a user's emotions and adjust the way they express the translation results based on that estimation. For example, if the user is relaxed, the system can provide a detailed translation. If the user is stressed, it can provide a concise translation. Furthermore, if the user is excited, it can provide a visually stimulating translation. This allows the system to provide translation results in an appropriate way that matches the user's emotions, thereby improving user satisfaction.

[0113] AI translation systems can estimate a user's emotions and adjust how the translation results are displayed based on that estimation. For example, if the user is stressed, the system can provide a simple and easy-to-read display. If the user is relaxed, the system can provide a display that includes more detailed information. Furthermore, if the user is in a hurry, the system can provide a concise and to-the-point display. This allows the system to provide translation results in a display format appropriate to the user's emotions, thereby improving user satisfaction.

[0114] AI translation systems can estimate a user's emotions and adjust the length of the translation based on that estimation. For example, if the user is relaxed, the system can provide a detailed translation. If the user is stressed, it can provide a concise translation. Furthermore, if the user is in a hurry, it can provide a short, to-the-point translation. This allows for translations of appropriate length according to the user's emotions, thereby improving user satisfaction.

[0115] The following briefly describes the processing flow for example form 2.

[0116] Step 1: The reception area receives user input. User input includes text input, voice input, and image input. The reception area is equipped with a keyboard or touchscreen for receiving text input, a microphone and voice recognition technology for receiving voice input, and a camera and image recognition technology for receiving image input. Step 2: The learning unit learns based on data from social media and television. The learning unit collects social media posts and television subtitle data, and uses machine learning algorithms and deep learning algorithms to learn the latest words and expressions. For example, it uses natural language processing techniques to analyze the meaning and context of words. Step 3: The analysis unit analyzes the input words. The analysis unit uses natural language processing technology to analyze the grammar and meaning of the input words, performing grammatical and semantic analysis. For example, it identifies the subject, predicate, and object of a sentence and understands the meaning of the input words. Step 4: The service provider provides the translation results. The service provider provides the translation results in text or audio format and is equipped with a display and speakers to show them to the user. For example, the translation results may be displayed on the display in text format and output from the speakers in audio format.

[0117] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0118] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0119] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0120] For example, the reception unit is implemented by the reception device 38 of the smart device 14. For example, the learning unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the provision unit is implemented by the output device 40 of the smart device 14. For example, the update unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the voice reception unit is implemented by the microphone 38B of the smart device 14. The correspondence between each unit and the devices and control units is not limited to the examples described above, and various changes are possible.

[0121] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0122] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0123] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0124] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0125] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0126] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0127] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0128] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0129] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0130] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0131] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0132] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0133] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0134] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0135] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0136] For example, the reception unit is implemented by the microphone 238 of the smart glasses 214. For example, the learning unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the provision unit is implemented by the speaker 240 of the smart glasses 214. For example, the update unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the voice reception unit is implemented by the microphone 238 of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the examples described above, and various changes are possible.

[0137] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0138] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0139] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0140] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0141] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0142] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0143] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0144] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0145] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0146] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0147] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0148] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0149] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0150] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0151] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0152] For example, the reception unit is implemented by the microphone 238 of the headset terminal 314. For example, the learning unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the provision unit is implemented by the speaker 240 of the headset terminal 314. For example, the update unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the voice reception unit is implemented by the microphone 238 of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the examples described above, and various changes are possible.

[0153] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0154] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0155] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0156] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0157] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0158] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0159] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0160] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0161] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0162] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0163] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0164] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0165] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0166] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0167] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0168] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0169] For example, the reception unit is implemented by the microphone 238 of the robot 414. For example, the learning unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the provision unit is implemented by the speaker 240 of the robot 414. For example, the update unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the voice reception unit is implemented by the microphone 238 of the robot 414. The correspondence between each unit and the device or control unit is not limited to the examples described above, and various changes are possible.

[0170] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0171] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0172] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0173] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0174] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0175] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0176] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0177] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0178] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0179] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0180] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0181] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0182] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0183] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0184] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0185] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0186] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0187] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0188] (Note 1) A reception area that receives user input, The learning section uses data from social media or television to guide students, An analysis unit that analyzes the input words, It includes a section that provides translation results. A system characterized by the following features. (Note 2) It includes an update unit for updating the training data. The system described in Appendix 1, characterized by the features described herein. (Note 3) It is equipped with a voice input receiving unit. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned reception unit is This provides a method for estimating user emotions and adjusting the input acceptance interface based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned reception unit is It analyzes the user's past input history and suggests appropriate input methods. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned reception unit is When receiving input, the system customizes the appropriate input method based on the user's current situation and environment. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned reception unit is It estimates the user's emotions and appropriately determines the priority of input content based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned reception unit is When receiving input, the system prioritizes accepting inputs that are highly relevant, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned reception unit is When receiving input, the system analyzes the user's social media activity and prioritizes accepting relevant input. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned learning unit, The system estimates the user's emotions and selects training data based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned learning unit, During training, the learning algorithm is optimized by referring to past training data. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned learning unit, During training, improve learning accuracy by using datasets specific to particular languages ​​and cultures. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned learning unit, It estimates the user's emotions and adjusts the learning frequency based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned learning unit, During learning, in addition to data from social media and television, data from blogs and news sites is added to the learning process. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned learning unit, During learning, the learning content is updated in real time by incorporating user feedback. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, The system estimates the user's emotions and adjusts the analysis algorithm based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, During analysis, the context of the input words is taken into consideration to improve analysis accuracy. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned analysis unit, During analysis, different analysis methods are applied depending on the category of the input words. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned analysis unit, It estimates the user's emotions and adjusts how the analysis results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned analysis unit, During analysis, the geographical context of the input words is taken into consideration. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned analysis unit, During analysis, the system improves analysis accuracy by referencing relevant literature for the input terms. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned supply unit is, It estimates the user's emotions and adjusts the way the translation results are expressed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned supply unit is, When providing the translation, adjust the level of detail displayed based on the importance of the translation result. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned supply unit is, When providing the translation results, different display algorithms are applied depending on the category of the translation. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned supply unit is, It estimates the user's emotions and adjusts the length of the translation result based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned supply unit is, When providing the translation, the display priority will be determined based on when the translation results were submitted. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned supply unit is, When providing the translations, the display order will be adjusted based on the relevance of the translation results. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned update section is, We estimate user sentiment and adjust the timing of updates based on that estimated sentiment. The system described in Appendix 2, characterized by the features described herein. (Note 29) The aforementioned update section is, During updates, the system will refer to past update history to select the most suitable update method. The system described in Appendix 2, characterized by the features described herein. (Note 30) The aforementioned update section is, It estimates user sentiment and determines update priorities based on the estimated user sentiment. The system described in Appendix 2, characterized by the features described herein. (Note 31) The aforementioned update section is, During the update, in addition to data from social media and television, data from blogs and news sites will also be added. The system described in Appendix 2, characterized by the features described herein. (Note 32) The aforementioned voice reception unit is It estimates the user's emotions and adjusts the voice input interface based on those emotions. The system described in Appendix 3, characterized by the features described herein. (Note 33) The aforementioned voice reception unit is When using voice input, the system refers to the user's past voice input history to suggest the optimal input method. The system described in Appendix 3, characterized by the features described herein. (Note 34) The aforementioned voice reception unit is It estimates the user's emotions and prioritizes voice input content based on the estimated emotions. The system described in Appendix 3, characterized by the features described herein. (Note 35) The aforementioned voice reception unit is When using voice input, the system prioritizes accepting voice input that is highly relevant, taking into account the user's geographical location. The system described in Appendix 3, characterized by the features described herein. [Explanation of Symbols]

[0189] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A reception area that receives user input, The learning section uses data from social media or television to guide students, An analysis unit that analyzes the input words, It includes a section that provides translation results. A system characterized by the following features.

2. It includes an update unit for updating the training data. The system according to feature 1.

3. It is equipped with a voice input receiving unit. The system according to feature 1.

4. The aforementioned reception unit is This provides a method for estimating user emotions and adjusting the input acceptance interface based on those estimated emotions. The system according to feature 1.

5. The aforementioned reception unit is It analyzes the user's past input history and suggests appropriate input methods. The system according to feature 1.

6. The aforementioned reception unit is When receiving input, the system customizes the appropriate input method based on the user's current situation and environment. The system according to feature 1.

7. The aforementioned reception unit is It estimates the user's emotions and appropriately determines the priority of input content based on the estimated user emotions. The system according to feature 1.

8. The aforementioned reception unit is When receiving input, the system prioritizes accepting inputs that are highly relevant, taking into account the user's geographical location. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A