system

The system addresses the challenge of providing real-time responses by using AI to collect and analyze data on a specific person, generating video and audio that replicates their knowledge and experience, enabling direct learning from their insights.

JP7869836B2Active Publication Date: 2026-06-03SOFTBANK GROUP CORP

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-09-19
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

Conventional systems struggle to provide real-time responses based on the knowledge and experience of a specific person.

Method used

A system comprising a collection unit, an analysis unit, and a generation unit that collects data on a specific person, analyzes user questions, and generates responses as video or audio using AI, employing technologies like deepfake and speech synthesis.

Benefits of technology

Enables real-time generation of responses that mimic the knowledge and experience of a specific person, allowing users to learn directly from their insights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007869836000001
    Figure 0007869836000001
  • Figure 0007869836000002
    Figure 0007869836000002
  • Figure 0007869836000003
    Figure 0007869836000003
Patent Text Reader

Abstract

To provide a system that provides an answer on the basis of the knowledge and experience of a specific person in real time.SOLUTION: A system includes a collection section, an analysis section, and a generation section. The collection section collects data on the specific person to cause the AI to learn. The analysis section analyzes a question or consultation from a user on the basis of the data collected by the collection section. The generation section outputs the answer generated by the analysis section as a video or a voice.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, there is a problem that it is difficult to provide a response based on the knowledge and experience of a specific person in real time.

[0005] The system according to the embodiment aims to provide a response based on the knowledge and experience of a specific person in real time.

Means for Solving the Problems

[0006] The system according to the embodiment includes a collection unit, an analysis unit, and a generation unit. The collection unit collects data of a specific person and makes it learned by AI. The analysis unit analyzes questions and consultations from users based on the data collected by the collection unit and generates responses. The generation unit outputs the responses generated by the analysis unit as video or audio. [Effects of the Invention]

[0007] The system according to the embodiment can provide answers in real time based on the knowledge and experience of a specific person. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The system according to an embodiment of the present invention is a system for developing an AI that has been trained on the words, actions, behavioral guidelines, life philosophy, and past publications of a specific person (for example, a famous business executive). This system can generate video and audio that sounds as if the specific person is actually speaking in response to questions and consultations from users. For example, the system first collects data such as the words, actions, behavioral guidelines, life philosophy, and past publications of a specific person and trains the AI ​​on it. In this process, the patterns of the specific person's statements and actions are analyzed in detail so that the AI ​​can imitate them. For example, by collecting interview videos or lecture recordings of a specific person and training the AI ​​on them, the way of speaking and expression of that person can be reproduced. Next, the system accepts questions and consultations from users as input. For example, a question such as "What is the secret to business success?" is entered. This input is given to the AI. The AI ​​analyzes the entered question or consultation and generates an appropriate answer based on the words, actions, and behavioral guidelines of a specific person. For example, it generates an answer about the secret to business success by referring to the content of the specific person's past statements and books. The generated answer is output as video and audio that sounds as if the specific person is actually speaking. For example, based on answers generated by AI, a video of a specific person is created, and this video then responds to the user. By replicating the specific person's way of speaking and expression, the user can experience what it's like to actually converse with that person. This mechanism allows users to directly learn from the knowledge and experience of a specific person and receive business and life advice. For instance, a young entrepreneur can receive advice from a specific person to gain concrete guidance for business success. Ordinary people can also learn from a specific person's life philosophy and behavioral guidelines to gain hints for self-improvement and growth. In this way, the system can generate answers to user questions and consultations using video and audio that mimics the specific person.

[0029] The system according to this embodiment comprises a collection unit, an analysis unit, and a generation unit. The collection unit collects data on a specific person and uses it to train an AI. The data on a specific person includes, but is not limited to, interview videos, lecture recordings, and book contents. For example, the collection unit can collect interview videos and use them to train the AI. The collection unit can also collect lecture recordings and use them to train the AI. Furthermore, the collection unit can collect book contents and use them to train the AI. For example, the collection unit can collect interview videos in high resolution and use them to train the AI. Lecture recordings can be collected as audio data and used to train the AI. Book contents can be collected as text data and used to train the AI. The analysis unit analyzes user questions and consultations based on the data collected by the collection unit and generates appropriate answers. For example, the analysis unit can analyze user questions using natural language processing technology. The analysis unit can also analyze questions using machine learning algorithms. Furthermore, the analysis unit can perform analysis based on past question data. For example, the analysis unit uses natural language processing technology to perform morphological analysis on the user's question and analyze its meaning. A machine learning algorithm learns from past question data and generates an appropriate answer to a new question. The generation unit outputs the answer generated by the analysis unit as video or audio. The generation unit generates video using, for example, deepfake technology. The generation unit can also generate audio using speech synthesis technology. Furthermore, the generation unit can generate video and audio in real time. For example, the generation unit generates video of a specific person using deepfake technology. Speech synthesis technology generates audio by mimicking the voice of a specific person. Real-time generation allows for the instantaneous generation of video and audio in response to a user's question. As a result, the system according to this embodiment can generate answers to user questions and consultations with video and audio that resemble those of a specific person.

[0030] The data collection unit collects data on specific individuals and uses it to train an AI. This data may include, but is not limited to, interview videos, lecture recordings, and book content. For example, the unit can collect interview videos and use them to train the AI. It can also collect lecture recordings and use them to train the AI. Furthermore, it can collect book content and use it to train the AI. For instance, the unit can collect interview videos in high resolution and use them to train the AI. Lecture recordings can be collected as audio data and used to train the AI. Book content can be collected as text data and used to train the AI. The data collection unit employs various technologies to efficiently collect this data. For example, interview videos can be collected in high resolution using video analysis technology, and important scenes within the video can be automatically extracted. Lecture recordings can be converted into text using speech recognition technology to facilitate content analysis. Book content can be extracted from printed books using OCR (optical character recognition) technology and digitized. This allows the data collection unit to centrally manage diverse data formats and efficiently train the AI. Furthermore, the data collection unit preprocesses the collected data to ensure data quality. For example, it performs noise reduction and resolution enhancement on video data, noise reduction and volume adjustment on audio data, and correction of typos and grammatical errors in text data. This improves the quality of the data the AI ​​learns from, enabling more accurate analysis and generation. Additionally, the data collection unit can adjust the frequency and scope of data collection, allowing for flexible responses to specific situations and conditions. For example, during specific events or lectures, data can be collected in real time and immediately used to train the AI. This enables the data collection unit to collect data efficiently and effectively, improving the overall system performance.

[0031] The analysis unit analyzes user questions and inquiries based on data collected by the data collection unit and generates appropriate answers. For example, the analysis unit uses natural language processing technology to analyze user questions. It can also analyze questions using machine learning algorithms. Furthermore, the analysis unit can perform analysis based on past question data. For instance, it uses natural language processing technology to perform morphological analysis on user questions and analyze their meaning. Machine learning algorithms learn from past question data and generate appropriate answers to new questions. By combining these technologies, the analysis unit achieves higher accuracy. Specifically, it uses natural language processing technology to perform morphological analysis on user questions to understand their context and intent. Next, it uses machine learning algorithms to derive the optimal answer by comparing it with past question data. Furthermore, the analysis unit continuously learns to improve the accuracy of its answers to user questions. For example, by collecting new question and answer data and regularly updating the model, it can always perform analysis based on the latest information. The analysis unit can also use anomaly detection algorithms to detect unusual patterns and abnormal questions and take appropriate action. This allows the analysis unit to respond quickly and accurately to a wide range of user questions and inquiries. Furthermore, the analysis unit can collect user feedback and continuously improve the accuracy and effectiveness of its analysis results. For example, it can adjust the parameters of its analysis algorithm based on user feedback to generate more appropriate answers. In addition, the analysis unit can improve the accuracy of its analysis by combining multiple analysis methods. For example, by combining natural language processing technology and machine learning algorithms, it can more accurately understand the intent of user questions and generate appropriate answers. As a result, the analysis unit can respond quickly and accurately to user questions and inquiries, improving the overall reliability and effectiveness of the system.

[0032] The generation unit outputs the answers generated by the analysis unit as video and audio. The generation unit can generate video using, for example, deepfake technology. It can also generate audio using speech synthesis technology. Furthermore, the generation unit can generate video and audio in real time. For example, the generation unit can generate video of a specific person using deepfake technology. Speech synthesis technology generates audio by mimicking the voice of a specific person. Real-time generation allows for instant video and audio generation in response to user questions. The generation unit utilizes these technologies to provide users with realistic answers. Specifically, it uses deepfake technology to realistically reproduce the facial expressions and mouth movements of a specific person, providing users with natural-looking video. Speech synthesis technology learns the characteristics of a specific person's voice and reproduces natural intonation and phrasing, providing users with realistic audio. Furthermore, the generation unit has high-speed processing capabilities to instantly generate video and audio in response to user questions. For example, it can use a high-performance GPU to execute deepfake technology in real time and instantly generate video in response to user questions. Furthermore, the speech synthesis technology boasts high processing capabilities, enabling it to instantly generate speech in response to user questions. This allows the generation unit to provide users with quick and natural answers. In addition, the generation unit continuously improves its technology to enhance the quality of the generated video and audio. For example, it introduces new deepfake and speech synthesis technologies to improve the quality of the generated video and audio. It also adjusts the parameters of the generation algorithm based on user feedback to generate more natural-sounding video and audio. As a result, the generation unit can provide users with high-quality video and audio, improving the overall reliability and effectiveness of the system.

[0033] The data collection unit can collect interview video or lecture recordings of specific individuals. For example, the data collection unit can collect interview video of a specific individual. Interview video includes, but is not limited to, television interviews or online interviews. For example, the data collection unit can collect television interviews in high resolution and use them to train the AI. The data collection unit can also collect online interviews and use them to train the AI. Furthermore, the data collection unit can also collect lecture recordings of specific individuals. Lecture recordings include, but is not limited to, academic conference presentations or seminar presentations. For example, the data collection unit can collect academic conference presentations as audio data and use them to train the AI. The data collection unit can also collect seminar presentations and use them to train the AI. As a result, by collecting interview video and lecture recordings of specific individuals, the data collection unit enables the AI ​​to more accurately imitate the words and actions of those individuals.

[0034] The analysis unit can analyze user questions using natural language processing techniques. For example, the analysis unit analyzes user questions using natural language processing techniques. Natural language processing techniques include, but are not limited to, morphological analysis, grammatical analysis, and semantic analysis. For example, the analysis unit can analyze user questions using morphological analysis. The analysis unit can also analyze questions using grammatical analysis. Furthermore, the analysis unit can also analyze questions using semantic analysis. For example, the analysis unit uses morphological analysis to break down user questions into individual words and analyze their meaning. Grammatical analysis analyzes the grammatical structure of a question and understands its meaning. Semantic analysis analyzes the context of a question and generates an appropriate answer. Thus, by using natural language processing techniques, the analysis unit can accurately analyze user questions and generate appropriate answers.

[0035] The generation unit can generate video or audio in real time using deepfake technology or speech synthesis technology. For example, the generation unit can generate video using deepfake technology. Deepfake technology includes, but is not limited to, face synthesis and voice synthesis. For example, the generation unit can generate video of a specific person using face synthesis. The generation unit can also generate audio using speech synthesis technology. Speech synthesis technology includes, but is not limited to, text-to-speech synthesis and voice imitation. For example, the generation unit can generate the voice of a specific person using text-to-speech synthesis. The generation unit can also generate the voice of a specific person using voice imitation. Furthermore, the generation unit can generate video and audio in real time. For example, the generation unit can use deepfake technology to instantly generate video of a specific person in response to a user's question. Speech synthesis technology can instantly generate the voice of a specific person in response to a user's question. Thus, the generation unit can generate video and audio that resemble a specific person in real time by using deepfake technology or speech synthesis technology.

[0036] The data collection unit can analyze the background and context of a specific person's statements and determine the priority of the data to collect. For example, the data collection unit can analyze the background and context of a specific person's statements and determine the priority of the data to collect. The analysis of background and context includes, but is not limited to, preceding and succeeding statements and related topics. For example, the data collection unit can analyze preceding and succeeding statements and prioritize the collection of content spoken by a specific person in a specific business setting. The data collection unit can also analyze related topics and, if a specific person's statement is related to a specific time or event, analyze and collect that background information. Furthermore, if a specific person's statement is related to a specific theme, the data collection unit can prioritize the collection of data related to that theme. For example, the data collection unit can prioritize the collection of content spoken by a specific person in a specific business setting. Furthermore, if a specific person's statement is related to a specific time or event, the data collection unit can also analyze and collect that background information. Furthermore, if a specific person's statement is related to a specific theme, the data collection unit can prioritize the collection of data related to that theme. This allows the data collection unit to prioritize the collection of more important data by analyzing the background and context of the statements.

[0037] The data collection unit can adjust the level of detail of the data it collects based on the frequency and importance of statements made by a particular person. For example, the data collection unit can adjust the level of detail of the data it collects based on the frequency and importance of statements made by a particular person. Frequency of statements includes, but is not limited to, the number of occurrences and time intervals. For example, the data collection unit can collect detailed data on topics that a particular person frequently mentions. The data collection unit can also collect data that includes detailed background information on important statements and famous quotes made by a particular person. Furthermore, if a particular person's statements are considered important in a particular field, the data collection unit can also collect detailed data related to that field. For example, the data collection unit can collect detailed data on topics that a particular person frequently mentions. The data collection unit can also collect data that includes detailed background information on important statements and famous quotes made by a particular person. Furthermore, if a particular person's statements are considered important in a particular field, the data collection unit can also collect detailed data related to that field. This allows the data collection unit to collect more accurate data by adjusting the level of detail of the data based on the frequency and importance of statements.

[0038] The data collection unit can select data to collect by considering the geographical context of a particular person's statements. For example, the data collection unit selects data to collect by considering the geographical context of a particular person's statements. Geographical context includes, but is not limited to, regional culture and geographical conditions. For example, the data collection unit collects lectures and interviews given by a particular person in a specific region. Furthermore, if a particular person's statements relate to a specific country or region, the data collection unit can also collect data related to that region. In addition, if a particular person's statements relate to a specific cultural or social background, the data collection unit can also collect data containing that background information. This allows the data collection unit to collect more relevant data by considering geographical context.

[0039] The data collection unit can collect and compare statements from other prominent figures related to the statements of a particular individual. For example, the data collection unit can collect and compare statements from other prominent figures related to the statements of a particular individual. These other prominent figures include, but are not limited to, prominent figures in the same industry or experts in related fields. The data collection unit can, for example, collect and compare statements from prominent figures contemporary with a particular individual. Furthermore, the data collection unit can collect and compare statements from experts in fields related to the statements of a particular individual. In addition, the data collection unit can collect and compare statements from prominent figures who hold different viewpoints than those of a particular individual. This allows the data collection unit to provide a more multifaceted perspective by comparing and analyzing statements from other prominent figures.

[0040] The analysis unit can analyze background information of a question and apply analytical methods to generate more specific answers. For example, the analysis unit can analyze background information of a question and apply analytical methods to generate more specific answers. Background information includes, but is not limited to, past questions and related topics. For example, the analysis unit can analyze past questions, and if the question is related to a specific business scenario, it can analyze the background information to generate specific answers. Furthermore, the analysis unit can analyze related topics, and if the question is related to a specific time or event, it can analyze the background information to generate specific answers. Furthermore, the analysis unit can analyze past questions and, if a question is related to a specific theme, analyze the background information to generate specific answers. This allows the analysis unit to generate more specific and appropriate answers by analyzing the background information of the question.

[0041] The analysis unit can improve the accuracy of its answers by applying different analysis algorithms depending on the category of the question. For example, the analysis unit can apply different analysis algorithms depending on the category of the question to improve the accuracy of its answers. Question categories include, for example, technical questions and business-related questions, but are not limited to these examples. For example, the analysis unit can apply a technical analysis algorithm to technical questions. It can also apply an analysis algorithm specialized in business strategy to business-related questions. Furthermore, it can apply an analysis algorithm that includes philosophical insights to questions about life philosophy. In this way, the analysis unit improves the accuracy of its answers by applying analysis algorithms appropriate to the category of the question.

[0042] The analysis unit can determine the priority of analysis based on when the questions were submitted. For example, the analysis unit can determine the priority of analysis based on when the questions were submitted. This includes, but is not limited to, the submission date and time, and urgency. For example, the analysis unit can determine the priority of questions based on the submission date and time. The analysis unit can also determine the priority of questions based on urgency. Furthermore, the analysis unit can also determine the priority of questions related to specific events or periods. This allows the analysis unit to respond quickly to urgent questions by determining the priority of analysis based on when the questions were submitted.

[0043] The analysis unit can adjust the order of analysis based on the relevance of the questions. For example, the analysis unit adjusts the order of analysis based on the relevance of the questions. Relevance includes, but is not limited to, the degree of topic relevance and relevance to past questions. For example, the analysis unit adjusts the order of questions based on the degree of topic relevance. The analysis unit can also adjust the order of questions based on their relevance to past questions. Furthermore, the analysis unit can adjust the order of questions related to specific themes. This allows the analysis unit to provide consistent answers to related questions by adjusting the order of analysis based on the relevance of the questions.

[0044] The generation unit can analyze the tone and pace of a specific person's speech and reflect it in the generated video and audio. For example, the generation unit can analyze the tone and pace of a specific person's speech and reflect it in the generated video and audio. The analysis of tone and pace includes, but is not limited to, speech analysis and rhythm analysis. For example, the generation unit can use speech analysis to analyze the tone of a specific person's speech. The generation unit can also use rhythm analysis to analyze the pace of a specific person's speech. Furthermore, the generation unit can combine speech analysis and rhythm analysis to analyze the tone and pace of a specific person's speech. For example, if a specific person speaks in a calm tone, the generation unit can use speech analysis to generate video and audio that reproduces that tone. Furthermore, if a specific person speaks quickly, the generation unit can use rhythm analysis to generate video and audio that reproduces that pace. Furthermore, if a specific person speaks with emphasis, the generation unit can combine speech analysis and rhythm analysis to generate video and audio that reproduces that emphasis. This allows the generation unit to create more realistic video and audio by reflecting the tone and pace of a specific person's speech.

[0045] The generation unit can adjust the level of detail in the generated video and audio based on the context of a specific person's statement. For example, the generation unit adjusts the level of detail in the generated video and audio based on the context of a specific person's statement. Contextual analysis includes, but is not limited to, preceding and succeeding statements and related topics. For example, the generation unit analyzes preceding and succeeding statements, and if a specific person's statement includes a concrete example, it generates video and audio that reproduces that example in detail. The generation unit can also analyze related topics, and if a specific person's statement includes an abstract concept, it can generate video and audio that visually represents that concept. Furthermore, the generation unit can combine preceding and succeeding statements with related topics, and if a specific person's statement includes an emotional element, it can generate video and audio that emphasizes that emotion. For example, the generation unit analyzes preceding and succeeding statements, and if a specific person's statement includes a concrete example, it generates video and audio that reproduces that example in detail. The generation unit can also analyze related topics, and if a specific person's statement includes an abstract concept, it can generate video and audio that visually represents that concept. Furthermore, the generation unit can combine preceding and succeeding statements with related topics to generate video and audio that emphasizes emotions when a particular person's statement contains emotional elements. This allows the generation unit to produce more appropriate video and audio by adjusting the level of detail based on the context of the statement.

[0046] The generation unit can customize the generated video and audio by taking into account the geographical context of a specific person's statements. For example, the generation unit can customize the generated video and audio by taking into account the geographical context of a specific person's statements. Geographical context includes, but is not limited to, regional culture and geographical conditions. For example, the generation unit can generate video and audio that recreates a speech given by a specific person in a specific region. Furthermore, if a specific person's statements relate to a specific country or region, the generation unit can also generate video and audio that includes the context of that region. In addition, if a specific person's statements relate to a specific cultural or social background, the generation unit can also generate video and audio that includes that background information. This allows the generation unit to create more relevant video and audio by taking geographical context into consideration.

[0047] The generation unit can generate and compare statements from other prominent figures related to a specific person's statements. For example, the generation unit can generate and compare statements from other prominent figures related to a specific person's statements. These other prominent figures include, but are not limited to, prominent figures in the same industry or experts in related fields. For example, the generation unit can generate video and audio that reproduces and compares statements from prominent figures contemporary with a specific person. Furthermore, the generation unit can generate video and audio that reproduces and compares statements from experts in related fields to a specific person's statements. In addition, the generation unit can generate video and audio that reproduces and compares statements from prominent figures with different viewpoints than those of a specific person. This allows the generation unit to provide a more multifaceted perspective by displaying statements in comparison with those of other prominent figures.

[0048] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0049] The data collection unit can also customize the types of data it collects based on the user's interests. For example, if a user is interested in a particular business field, it can prioritize collecting statements and book contents from specific individuals related to that field. Similarly, if a user is interested in self-improvement, it can collect data on the life philosophy and guiding principles of specific individuals. Furthermore, if a user is seeking information related to a specific time period or event, it can collect statements and activities from specific individuals related to that time period or event. This allows the data collection unit to provide more personalized information by customizing data collection according to the user's interests.

[0050] The analytics unit can analyze the user's questioning trends based on their past question history and generate more appropriate answers. For example, if a user has asked many business-related questions in the past, it will prioritize generating business-related answers. Similarly, if a user has asked many self-improvement-related questions, it can generate self-improvement-related answers. Furthermore, if a user repeatedly asks questions on a specific theme, it can generate detailed answers related to that theme. In this way, the analytics unit can provide more appropriate and personalized answers by utilizing the user's question history.

[0051] The data collection unit can also customize the data it collects, taking into account the user's geographical background. For example, if a user lives in a specific region, it can prioritize collecting statements and activities of specific individuals associated with that region. Furthermore, if a user has a specific cultural or social background, it can collect data related to that background. Additionally, if a user is interested in a particular country or region, it can collect statements and activities of specific individuals associated with that country or region. This allows the data collection unit to provide more relevant information by customizing data collection according to the user's geographical background.

[0052] The generation unit can also generate videos and audio tailored to the user's preferences based on their past question history. For example, if a user has asked many business-related questions in the past, it will prioritize generating business-related videos and audio. Similarly, if a user has asked many self-improvement-related questions, it can generate videos and audio related to self-improvement. Furthermore, if a user repeatedly asks questions on a specific theme, it can generate detailed videos and audio related to that theme. In this way, the generation unit can provide more personalized videos and audio by utilizing the user's question history.

[0053] The analysis unit can also customize its analysis methods based on the category of the user's question. For example, technical questions can be analyzed using technical methods. Business-related questions can be analyzed using methods specifically tailored to business strategies. Furthermore, questions about life philosophy can be analyzed using methods that include philosophical insights. This allows the analysis unit to provide more accurate answers by applying analysis methods appropriate to the category of the question.

[0054] The data collection unit can also prioritize the data it collects based on the user's interests. For example, if a user is interested in a particular business field, it can prioritize collecting statements and book contents from specific individuals related to that field. Similarly, if a user is interested in self-improvement, it can prioritize collecting data on the life philosophy and behavioral guidelines of specific individuals. Furthermore, if a user is seeking information related to a specific time or event, it can prioritize collecting statements and activities from specific individuals related to that time or event. This allows the data collection unit to customize data collection according to the user's interests, thereby providing more personalized information.

[0055] The following briefly describes the processing flow for example form 1.

[0056] Step 1: The data collection unit collects data on a specific person (for example, a specific individual) and uses it to train the AI. The collected data includes interview videos, lecture recordings, and book contents. The data collection unit collects interview videos in high resolution and uses them to train the AI. Lecture recordings are collected as audio data and used to train the AI. Book contents are collected as text data and used to train the AI. Step 2: The analysis unit analyzes user questions and inquiries based on the data collected by the collection unit and generates appropriate answers. The analysis unit can analyze user questions using natural language processing technology and machine learning algorithms, and can also perform analysis based on past question data. For example, it can use natural language processing technology to perform morphological analysis on user questions and analyze their meaning. Machine learning algorithms learn from past question data and generate appropriate answers to new questions. Step 3: The generation unit outputs the answers generated by the analysis unit as video and audio. The generation unit can generate video using deepfake technology and audio using speech synthesis technology. Furthermore, the generation unit can generate video and audio in real time. For example, it can generate video of a specific person using deepfake technology and generate audio by mimicking the voice of a specific person using speech synthesis technology. Real-time generation allows for the instantaneous generation of video and audio in response to user questions.

[0057] (Example of form 2) The system according to an embodiment of the present invention is a system for developing an AI that has been trained on the words, actions, behavioral guidelines, life philosophy, and past publications of a specific person (for example, a famous business executive). This system can generate video and audio that sounds as if the specific person is actually speaking in response to questions and consultations from users. For example, the system first collects data such as the words, actions, behavioral guidelines, life philosophy, and past publications of a specific person and trains the AI ​​on it. In this process, the patterns of the specific person's statements and actions are analyzed in detail so that the AI ​​can imitate them. For example, by collecting interview videos or lecture recordings of a specific person and training the AI ​​on them, the way of speaking and expression of that person can be reproduced. Next, the system accepts questions and consultations from users as input. For example, a question such as "What is the secret to business success?" is entered. This input is given to the AI. The AI ​​analyzes the entered question or consultation and generates an appropriate answer based on the words, actions, and behavioral guidelines of a specific person. For example, it generates an answer about the secret to business success by referring to the content of the specific person's past statements and books. The generated answer is output as video and audio that sounds as if the specific person is actually speaking. For example, based on answers generated by AI, a video of a specific person is created, and this video then responds to the user. By replicating the specific person's way of speaking and expression, the user can experience what it's like to actually converse with that person. This mechanism allows users to directly learn from the knowledge and experience of a specific person and receive business and life advice. For instance, a young entrepreneur can receive advice from a specific person to gain concrete guidance for business success. Ordinary people can also learn from a specific person's life philosophy and behavioral guidelines to gain hints for self-improvement and growth. In this way, the system can generate answers to user questions and consultations using video and audio that mimics the specific person.

[0058] The system according to this embodiment comprises a collection unit, an analysis unit, and a generation unit. The collection unit collects data on a specific person and uses it to train an AI. The data on a specific person includes, but is not limited to, interview videos, lecture recordings, and book contents. For example, the collection unit can collect interview videos and use them to train the AI. The collection unit can also collect lecture recordings and use them to train the AI. Furthermore, the collection unit can collect book contents and use them to train the AI. For example, the collection unit can collect interview videos in high resolution and use them to train the AI. Lecture recordings can be collected as audio data and used to train the AI. Book contents can be collected as text data and used to train the AI. The analysis unit analyzes user questions and consultations based on the data collected by the collection unit and generates appropriate answers. For example, the analysis unit can analyze user questions using natural language processing technology. The analysis unit can also analyze questions using machine learning algorithms. Furthermore, the analysis unit can perform analysis based on past question data. For example, the analysis unit uses natural language processing technology to perform morphological analysis on the user's question and analyze its meaning. A machine learning algorithm learns from past question data and generates an appropriate answer to a new question. The generation unit outputs the answer generated by the analysis unit as video or audio. The generation unit generates video using, for example, deepfake technology. The generation unit can also generate audio using speech synthesis technology. Furthermore, the generation unit can generate video and audio in real time. For example, the generation unit generates video of a specific person using deepfake technology. Speech synthesis technology generates audio by mimicking the voice of a specific person. Real-time generation allows for the instantaneous generation of video and audio in response to a user's question. As a result, the system according to this embodiment can generate answers to user questions and consultations with video and audio that resemble those of a specific person.

[0059] The data collection unit collects data on specific individuals and uses it to train an AI. This data may include, but is not limited to, interview videos, lecture recordings, and book content. For example, the unit can collect interview videos and use them to train the AI. It can also collect lecture recordings and use them to train the AI. Furthermore, it can collect book content and use it to train the AI. For instance, the unit can collect interview videos in high resolution and use them to train the AI. Lecture recordings can be collected as audio data and used to train the AI. Book content can be collected as text data and used to train the AI. The data collection unit employs various technologies to efficiently collect this data. For example, interview videos can be collected in high resolution using video analysis technology, and important scenes within the video can be automatically extracted. Lecture recordings can be converted into text using speech recognition technology to facilitate content analysis. Book content can be extracted from printed books using OCR (optical character recognition) technology and digitized. This allows the data collection unit to centrally manage diverse data formats and efficiently train the AI. Furthermore, the data collection unit preprocesses the collected data to ensure data quality. For example, it performs noise reduction and resolution enhancement on video data, noise reduction and volume adjustment on audio data, and correction of typos and grammatical errors in text data. This improves the quality of the data the AI ​​learns from, enabling more accurate analysis and generation. Additionally, the data collection unit can adjust the frequency and scope of data collection, allowing for flexible responses to specific situations and conditions. For example, during specific events or lectures, data can be collected in real time and immediately used to train the AI. This enables the data collection unit to collect data efficiently and effectively, improving the overall system performance.

[0060] The analysis unit analyzes user questions and inquiries based on data collected by the data collection unit and generates appropriate answers. For example, the analysis unit uses natural language processing technology to analyze user questions. It can also analyze questions using machine learning algorithms. Furthermore, the analysis unit can perform analysis based on past question data. For instance, it uses natural language processing technology to perform morphological analysis on user questions and analyze their meaning. Machine learning algorithms learn from past question data and generate appropriate answers to new questions. By combining these technologies, the analysis unit achieves higher accuracy. Specifically, it uses natural language processing technology to perform morphological analysis on user questions to understand their context and intent. Next, it uses machine learning algorithms to derive the optimal answer by comparing it with past question data. Furthermore, the analysis unit continuously learns to improve the accuracy of its answers to user questions. For example, by collecting new question and answer data and regularly updating the model, it can always perform analysis based on the latest information. The analysis unit can also use anomaly detection algorithms to detect unusual patterns and abnormal questions and take appropriate action. This allows the analysis unit to respond quickly and accurately to a wide range of user questions and inquiries. Furthermore, the analysis unit can collect user feedback and continuously improve the accuracy and effectiveness of its analysis results. For example, it can adjust the parameters of its analysis algorithm based on user feedback to generate more appropriate answers. In addition, the analysis unit can improve the accuracy of its analysis by combining multiple analysis methods. For example, by combining natural language processing technology and machine learning algorithms, it can more accurately understand the intent of user questions and generate appropriate answers. As a result, the analysis unit can respond quickly and accurately to user questions and inquiries, improving the overall reliability and effectiveness of the system.

[0061] The generation unit outputs the answers generated by the analysis unit as video and audio. The generation unit can generate video using, for example, deepfake technology. It can also generate audio using speech synthesis technology. Furthermore, the generation unit can generate video and audio in real time. For example, the generation unit can generate video of a specific person using deepfake technology. Speech synthesis technology generates audio by mimicking the voice of a specific person. Real-time generation allows for instant video and audio generation in response to user questions. The generation unit utilizes these technologies to provide users with realistic answers. Specifically, it uses deepfake technology to realistically reproduce the facial expressions and mouth movements of a specific person, providing users with natural-looking video. Speech synthesis technology learns the characteristics of a specific person's voice and reproduces natural intonation and phrasing, providing users with realistic audio. Furthermore, the generation unit has high-speed processing capabilities to instantly generate video and audio in response to user questions. For example, it can use a high-performance GPU to execute deepfake technology in real time and instantly generate video in response to user questions. Furthermore, the speech synthesis technology boasts high processing capabilities, enabling it to instantly generate speech in response to user questions. This allows the generation unit to provide users with quick and natural answers. In addition, the generation unit continuously improves its technology to enhance the quality of the generated video and audio. For example, it introduces new deepfake and speech synthesis technologies to improve the quality of the generated video and audio. It also adjusts the parameters of the generation algorithm based on user feedback to generate more natural-sounding video and audio. As a result, the generation unit can provide users with high-quality video and audio, improving the overall reliability and effectiveness of the system.

[0062] The data collection unit can collect interview video or lecture recordings of specific individuals. For example, the data collection unit can collect interview video of a specific individual. Interview video includes, but is not limited to, television interviews or online interviews. For example, the data collection unit can collect television interviews in high resolution and use them to train the AI. The data collection unit can also collect online interviews and use them to train the AI. Furthermore, the data collection unit can also collect lecture recordings of specific individuals. Lecture recordings include, but is not limited to, academic conference presentations or seminar presentations. For example, the data collection unit can collect academic conference presentations as audio data and use them to train the AI. The data collection unit can also collect seminar presentations and use them to train the AI. As a result, by collecting interview video and lecture recordings of specific individuals, the data collection unit enables the AI ​​to more accurately imitate the words and actions of those individuals.

[0063] The analysis unit can analyze user questions using natural language processing techniques. For example, the analysis unit analyzes user questions using natural language processing techniques. Natural language processing techniques include, but are not limited to, morphological analysis, grammatical analysis, and semantic analysis. For example, the analysis unit can analyze user questions using morphological analysis. The analysis unit can also analyze questions using grammatical analysis. Furthermore, the analysis unit can also analyze questions using semantic analysis. For example, the analysis unit uses morphological analysis to break down user questions into individual words and analyze their meaning. Grammatical analysis analyzes the grammatical structure of a question and understands its meaning. Semantic analysis analyzes the context of a question and generates an appropriate answer. Thus, by using natural language processing techniques, the analysis unit can accurately analyze user questions and generate appropriate answers.

[0064] The generation unit can generate video or audio in real time using deepfake technology or speech synthesis technology. For example, the generation unit can generate video using deepfake technology. Deepfake technology includes, but is not limited to, face synthesis and voice synthesis. For example, the generation unit can generate video of a specific person using face synthesis. The generation unit can also generate audio using speech synthesis technology. Speech synthesis technology includes, but is not limited to, text-to-speech synthesis and voice imitation. For example, the generation unit can generate the voice of a specific person using text-to-speech synthesis. The generation unit can also generate the voice of a specific person using voice imitation. Furthermore, the generation unit can generate video and audio in real time. For example, the generation unit can use deepfake technology to instantly generate video of a specific person in response to a user's question. Speech synthesis technology can instantly generate the voice of a specific person in response to a user's question. Thus, the generation unit can generate video and audio that resemble a specific person in real time by using deepfake technology or speech synthesis technology.

[0065] The data collection unit can estimate the user's emotions and adjust the types of data it collects based on those estimated emotions. For example, the data collection unit can estimate the user's emotions and adjust the types of data it collects based on those estimated emotions. Emotion estimation includes, but is not limited to, facial expression analysis, voice analysis, and text analysis. For example, the data collection unit can use facial expression analysis to estimate the user's emotions. It can also use voice analysis to estimate the user's emotions. Furthermore, it can use text analysis to estimate the user's emotions. For example, if the user is feeling stressed, the data collection unit can prioritize collecting interview videos or lecture recordings of specific individuals that are relaxing. It can also use voice analysis to collect data containing success stories or encouraging words if the user wants to boost their motivation. Furthermore, it can use text analysis to collect data containing business strategies or practical advice if the user is seeking specific advice. This allows the data collection unit to collect more relevant data by adjusting the types of data it collects according to the user's emotions.

[0066] The data collection unit can analyze the background and context of a specific person's statements and determine the priority of the data to collect. For example, the data collection unit can analyze the background and context of a specific person's statements and determine the priority of the data to collect. The analysis of background and context includes, but is not limited to, preceding and succeeding statements and related topics. For example, the data collection unit can analyze preceding and succeeding statements and prioritize the collection of content spoken by a specific person in a specific business setting. The data collection unit can also analyze related topics and, if a specific person's statement is related to a specific time or event, analyze and collect that background information. Furthermore, if a specific person's statement is related to a specific theme, the data collection unit can prioritize the collection of data related to that theme. For example, the data collection unit can prioritize the collection of content spoken by a specific person in a specific business setting. Furthermore, if a specific person's statement is related to a specific time or event, the data collection unit can also analyze and collect that background information. Furthermore, if a specific person's statement is related to a specific theme, the data collection unit can prioritize the collection of data related to that theme. This allows the data collection unit to prioritize the collection of more important data by analyzing the background and context of the statements.

[0067] The data collection unit can adjust the level of detail of the data it collects based on the frequency and importance of statements made by a particular person. For example, the data collection unit can adjust the level of detail of the data it collects based on the frequency and importance of statements made by a particular person. Frequency of statements includes, but is not limited to, the number of occurrences and time intervals. For example, the data collection unit can collect detailed data on topics that a particular person frequently mentions. The data collection unit can also collect data that includes detailed background information on important statements and famous quotes made by a particular person. Furthermore, if a particular person's statements are considered important in a particular field, the data collection unit can also collect detailed data related to that field. For example, the data collection unit can collect detailed data on topics that a particular person frequently mentions. The data collection unit can also collect data that includes detailed background information on important statements and famous quotes made by a particular person. Furthermore, if a particular person's statements are considered important in a particular field, the data collection unit can also collect detailed data related to that field. This allows the data collection unit to collect more accurate data by adjusting the level of detail of the data based on the frequency and importance of statements.

[0068] The data collection unit can estimate the user's emotions and adjust the timing of data collection based on the estimated emotions. For example, the data collection unit can estimate the user's emotions and adjust the timing of data collection based on the estimated emotions. Emotion estimation includes, but is not limited to, facial expression analysis, voice analysis, and text analysis. For example, the data collection unit can estimate the user's emotions using facial expression analysis. It can also estimate the user's emotions using voice analysis. Furthermore, it can estimate the user's emotions using text analysis. For example, if the user is in a hurry, the data collection unit can prioritize collecting short, effective data. It can also use voice analysis to collect long lecture recordings or interview videos if the user is relaxed. Furthermore, it can use text analysis to collect data at a specific time if the user is seeking advice at that time. This allows the data collection unit to collect data at a more appropriate time by adjusting the timing of data collection according to the user's emotions.

[0069] The data collection unit can select data to collect by considering the geographical context of a particular person's statements. For example, the data collection unit selects data to collect by considering the geographical context of a particular person's statements. Geographical context includes, but is not limited to, regional culture and geographical conditions. For example, the data collection unit collects lectures and interviews given by a particular person in a specific region. Furthermore, if a particular person's statements relate to a specific country or region, the data collection unit can also collect data related to that region. In addition, if a particular person's statements relate to a specific cultural or social background, the data collection unit can also collect data containing that background information. This allows the data collection unit to collect more relevant data by considering geographical context.

[0070] The data collection unit can collect and compare statements from other prominent figures related to the statements of a particular individual. For example, the data collection unit can collect and compare statements from other prominent figures related to the statements of a particular individual. These other prominent figures include, but are not limited to, prominent figures in the same industry or experts in related fields. The data collection unit can, for example, collect and compare statements from prominent figures contemporary with a particular individual. Furthermore, the data collection unit can collect and compare statements from experts in fields related to the statements of a particular individual. In addition, the data collection unit can collect and compare statements from prominent figures who hold different viewpoints than those of a particular individual. This allows the data collection unit to provide a more multifaceted perspective by comparing and analyzing statements from other prominent figures.

[0071] The analysis unit can estimate the user's emotions and adjust the analysis algorithm based on the estimated emotions. For example, the analysis unit can estimate the user's emotions and adjust the analysis algorithm based on the estimated emotions. Emotion estimation includes, but is not limited to, facial expression analysis, voice analysis, and text analysis. For example, the analysis unit can estimate the user's emotions using facial expression analysis. It can also estimate the user's emotions using voice analysis. Furthermore, it can estimate the user's emotions using text analysis. For example, the analysis unit can use facial expression analysis to apply an algorithm that generates a concise and clear response if the user is stressed. It can also use voice analysis to apply an algorithm that generates a detailed and insightful response if the user is relaxed. Furthermore, it can use text analysis to apply an algorithm that generates a quick response if the user is in a hurry. This allows the analysis unit to generate more appropriate responses by adjusting the analysis algorithm according to the user's emotions.

[0072] The analysis unit can analyze background information of a question and apply analytical methods to generate more specific answers. For example, the analysis unit can analyze background information of a question and apply analytical methods to generate more specific answers. Background information includes, but is not limited to, past questions and related topics. For example, the analysis unit can analyze past questions, and if the question is related to a specific business scenario, it can analyze the background information to generate specific answers. Furthermore, the analysis unit can analyze related topics, and if the question is related to a specific time or event, it can analyze the background information to generate specific answers. Furthermore, the analysis unit can analyze past questions and, if a question is related to a specific theme, analyze the background information to generate specific answers. This allows the analysis unit to generate more specific and appropriate answers by analyzing the background information of the question.

[0073] The analysis unit can improve the accuracy of its answers by applying different analysis algorithms depending on the category of the question. For example, the analysis unit can apply different analysis algorithms depending on the category of the question to improve the accuracy of its answers. Question categories include, for example, technical questions and business-related questions, but are not limited to these examples. For example, the analysis unit can apply a technical analysis algorithm to technical questions. It can also apply an analysis algorithm specialized in business strategy to business-related questions. Furthermore, it can apply an analysis algorithm that includes philosophical insights to questions about life philosophy. In this way, the analysis unit improves the accuracy of its answers by applying analysis algorithms appropriate to the category of the question.

[0074] The analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated emotions. For example, the analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated emotions. Emotion estimation includes, but is not limited to, facial expression analysis, voice analysis, and text analysis. For example, the analysis unit can estimate the user's emotions using facial expression analysis. The analysis unit can also estimate the user's emotions using voice analysis. Furthermore, the analysis unit can also estimate the user's emotions using text analysis. For example, the analysis unit can use facial expression analysis to provide a simple and highly visible display method when the user is tense. The analysis unit can also use voice analysis to provide a display method that includes detailed information when the user is relaxed. Furthermore, the analysis unit can use text analysis to provide a concise display method when the user is in a hurry. In this way, the analysis unit can provide a more appropriate display by adjusting the display method of the analysis results according to the user's emotions.

[0075] The analysis unit can determine the priority of analysis based on when the questions were submitted. For example, the analysis unit can determine the priority of analysis based on when the questions were submitted. This includes, but is not limited to, the submission date and time, and urgency. For example, the analysis unit can determine the priority of questions based on the submission date and time. The analysis unit can also determine the priority of questions based on urgency. Furthermore, the analysis unit can also determine the priority of questions related to specific events or periods. This allows the analysis unit to respond quickly to urgent questions by determining the priority of analysis based on when the questions were submitted.

[0076] The analysis unit can adjust the order of analysis based on the relevance of the questions. For example, the analysis unit adjusts the order of analysis based on the relevance of the questions. Relevance includes, but is not limited to, the degree of topic relevance and relevance to past questions. For example, the analysis unit adjusts the order of questions based on the degree of topic relevance. The analysis unit can also adjust the order of questions based on their relevance to past questions. Furthermore, the analysis unit can adjust the order of questions related to specific themes. This allows the analysis unit to provide consistent answers to related questions by adjusting the order of analysis based on the relevance of the questions.

[0077] The generation unit can estimate the user's emotions and adjust the way it expresses the generated video and audio based on those estimated emotions. For example, the generation unit can estimate the user's emotions and adjust the way it expresses the generated video and audio based on those estimated emotions. Emotion estimation includes, but is not limited to, facial expression analysis, voice analysis, and text analysis. For example, the generation unit can estimate the user's emotions using facial expression analysis. The generation unit can also estimate the user's emotions using voice analysis. Furthermore, the generation unit can also estimate the user's emotions using text analysis. For example, if the user is relaxed, the generation unit can generate video and audio with a calm tone of voice using facial expression analysis. If the user is in a hurry, the generation unit can also generate video and audio with a quick and concise expression using voice analysis. Furthermore, if the user is excited, the generation unit can generate video and audio with visually stimulating effects using text analysis. In this way, the generation unit can adjust the way it expresses the video and audio according to the user's emotions, enabling more appropriate expression.

[0078] The generation unit can analyze the tone and pace of a specific person's speech and reflect it in the generated video and audio. For example, the generation unit can analyze the tone and pace of a specific person's speech and reflect it in the generated video and audio. The analysis of tone and pace includes, but is not limited to, speech analysis and rhythm analysis. For example, the generation unit can use speech analysis to analyze the tone of a specific person's speech. The generation unit can also use rhythm analysis to analyze the pace of a specific person's speech. Furthermore, the generation unit can combine speech analysis and rhythm analysis to analyze the tone and pace of a specific person's speech. For example, if a specific person speaks in a calm tone, the generation unit can use speech analysis to generate video and audio that reproduces that tone. Furthermore, if a specific person speaks quickly, the generation unit can use rhythm analysis to generate video and audio that reproduces that pace. Furthermore, if a specific person speaks with emphasis, the generation unit can combine speech analysis and rhythm analysis to generate video and audio that reproduces that emphasis. This allows the generation unit to create more realistic video and audio by reflecting the tone and pace of a specific person's speech.

[0079] The generation unit can adjust the level of detail in the generated video and audio based on the context of a specific person's statement. For example, the generation unit adjusts the level of detail in the generated video and audio based on the context of a specific person's statement. Contextual analysis includes, but is not limited to, preceding and succeeding statements and related topics. For example, the generation unit analyzes preceding and succeeding statements, and if a specific person's statement includes a concrete example, it generates video and audio that reproduces that example in detail. The generation unit can also analyze related topics, and if a specific person's statement includes an abstract concept, it can generate video and audio that visually represents that concept. Furthermore, the generation unit can combine preceding and succeeding statements with related topics, and if a specific person's statement includes an emotional element, it can generate video and audio that emphasizes that emotion. For example, the generation unit analyzes preceding and succeeding statements, and if a specific person's statement includes a concrete example, it generates video and audio that reproduces that example in detail. The generation unit can also analyze related topics, and if a specific person's statement includes an abstract concept, it can generate video and audio that visually represents that concept. Furthermore, the generation unit can combine preceding and succeeding statements with related topics to generate video and audio that emphasizes emotions when a particular person's statement contains emotional elements. This allows the generation unit to produce more appropriate video and audio by adjusting the level of detail based on the context of the statement.

[0080] The generation unit can estimate the user's emotions and adjust the length of the generated video and audio based on the estimated emotions. For example, the generation unit can estimate the user's emotions and adjust the length of the generated video and audio based on the estimated emotions. Emotion estimation includes, but is not limited to, facial expression analysis, voice analysis, and text analysis. For example, the generation unit can estimate the user's emotions using facial expression analysis. The generation unit can also estimate the user's emotions using voice analysis. Furthermore, the generation unit can also estimate the user's emotions using text analysis. For example, the generation unit can use facial expression analysis to generate short, concise video and audio if the user is in a hurry. The generation unit can also use voice analysis to generate longer video and audio with detailed explanations if the user is relaxed. Furthermore, the generation unit can use text analysis to generate video and audio with visually stimulating effects if the user is excited. In this way, the generation unit can generate video and audio of a more appropriate length by adjusting the length according to the user's emotions.

[0081] The generation unit can customize the generated video and audio by taking into account the geographical context of a specific person's statements. For example, the generation unit can customize the generated video and audio by taking into account the geographical context of a specific person's statements. Geographical context includes, but is not limited to, regional culture and geographical conditions. For example, the generation unit can generate video and audio that recreates a speech given by a specific person in a specific region. Furthermore, if a specific person's statements relate to a specific country or region, the generation unit can also generate video and audio that includes the context of that region. In addition, if a specific person's statements relate to a specific cultural or social background, the generation unit can also generate video and audio that includes that background information. This allows the generation unit to create more relevant video and audio by taking geographical context into consideration.

[0082] The generation unit can generate and compare statements from other prominent figures related to a specific person's statements. For example, the generation unit can generate and compare statements from other prominent figures related to a specific person's statements. These other prominent figures include, but are not limited to, prominent figures in the same industry or experts in related fields. For example, the generation unit can generate video and audio that reproduces and compares statements from prominent figures contemporary with a specific person. Furthermore, the generation unit can generate video and audio that reproduces and compares statements from experts in related fields to a specific person's statements. In addition, the generation unit can generate video and audio that reproduces and compares statements from prominent figures with different viewpoints than those of a specific person. This allows the generation unit to provide a more multifaceted perspective by displaying statements in comparison with those of other prominent figures.

[0083] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0084] The data collection unit can also customize the types of data it collects based on the user's interests. For example, if a user is interested in a particular business field, it can prioritize collecting statements and book contents from specific individuals related to that field. Similarly, if a user is interested in self-improvement, it can collect data on the life philosophy and guiding principles of specific individuals. Furthermore, if a user is seeking information related to a specific time period or event, it can collect statements and activities from specific individuals related to that time period or event. This allows the data collection unit to provide more personalized information by customizing data collection according to the user's interests.

[0085] The analytics unit can analyze the user's questioning trends based on their past question history and generate more appropriate answers. For example, if a user has asked many business-related questions in the past, it will prioritize generating business-related answers. Similarly, if a user has asked many self-improvement-related questions, it can generate self-improvement-related answers. Furthermore, if a user repeatedly asks questions on a specific theme, it can generate detailed answers related to that theme. In this way, the analytics unit can provide more appropriate and personalized answers by utilizing the user's question history.

[0086] The generation unit can also estimate the user's emotions and adjust the tone and tempo of the generated video and audio based on those estimated emotions. For example, if the user is relaxed, it can generate video and audio with a calm tone and slow speech. If the user is in a hurry, it can generate video and audio with a quick and concise tone. Furthermore, if the user is excited, it can generate video and audio with an energetic tone. In this way, the generation unit can adjust the tone and tempo of the video and audio according to the user's emotions, enabling more appropriate expression.

[0087] The data collection unit can also customize the data it collects, taking into account the user's geographical background. For example, if a user lives in a specific region, it can prioritize collecting statements and activities of specific individuals associated with that region. Furthermore, if a user has a specific cultural or social background, it can collect data related to that background. Additionally, if a user is interested in a particular country or region, it can collect statements and activities of specific individuals associated with that country or region. This allows the data collection unit to provide more relevant information by customizing data collection according to the user's geographical background.

[0088] The analysis unit can estimate the user's emotions and adjust the analysis priority based on those emotions. For example, if the user is feeling stressed, it will prioritize generating responses that help reduce stress. Similarly, if the user wants to increase their motivation, it can prioritize generating responses that offer encouragement or share success stories. Furthermore, if the user is seeking specific advice, it can prioritize generating responses that include practical advice. This allows the analysis unit to provide more appropriate responses by adjusting the analysis priority according to the user's emotions.

[0089] The generation unit can also generate videos and audio tailored to the user's preferences based on their past question history. For example, if a user has asked many business-related questions in the past, it will prioritize generating business-related videos and audio. Similarly, if a user has asked many self-improvement-related questions, it can generate videos and audio related to self-improvement. Furthermore, if a user repeatedly asks questions on a specific theme, it can generate detailed videos and audio related to that theme. In this way, the generation unit can provide more personalized videos and audio by utilizing the user's question history.

[0090] The data collection unit can also estimate the user's emotions and adjust the amount of data collected based on that estimation. For example, if the user is stressed, it can collect a small amount of data and provide a concise response. Conversely, if the user is relaxed, it can collect a large amount of data and provide a detailed response. Furthermore, if the user is excited, it can collect visually stimulating data. In this way, the data collection unit can provide more relevant information by adjusting the amount of data according to the user's emotions.

[0091] The analysis unit can also customize its analysis methods based on the category of the user's question. For example, technical questions can be analyzed using technical methods. Business-related questions can be analyzed using methods specifically tailored to business strategies. Furthermore, questions about life philosophy can be analyzed using methods that include philosophical insights. This allows the analysis unit to provide more accurate answers by applying analysis methods appropriate to the category of the question.

[0092] The generation unit can also estimate the user's emotions and adjust the style of the generated video and audio based on those emotions. For example, if the user is relaxed, it can generate calm-style video and audio. If the user is in a hurry, it can generate fast-paced and concise-style video and audio. Furthermore, if the user is excited, it can generate visually stimulating-style video and audio. In this way, the generation unit can adjust the style of video and audio according to the user's emotions, enabling more appropriate expression.

[0093] The data collection unit can also prioritize the data it collects based on the user's interests. For example, if a user is interested in a particular business field, it can prioritize collecting statements and book contents from specific individuals related to that field. Similarly, if a user is interested in self-improvement, it can prioritize collecting data on the life philosophy and behavioral guidelines of specific individuals. Furthermore, if a user is seeking information related to a specific time or event, it can prioritize collecting statements and activities from specific individuals related to that time or event. This allows the data collection unit to customize data collection according to the user's interests, thereby providing more personalized information.

[0094] The following briefly describes the processing flow for example form 2.

[0095] Step 1: The data collection unit collects data on a specific person (for example, a specific individual) and uses it to train the AI. The collected data includes interview videos, lecture recordings, and book contents. The data collection unit collects interview videos in high resolution and uses them to train the AI. Lecture recordings are collected as audio data and used to train the AI. Book contents are collected as text data and used to train the AI. Step 2: The analysis unit analyzes user questions and inquiries based on the data collected by the collection unit and generates appropriate answers. The analysis unit can analyze user questions using natural language processing technology and machine learning algorithms, and can also perform analysis based on past question data. For example, it can use natural language processing technology to perform morphological analysis on user questions and analyze their meaning. Machine learning algorithms learn from past question data and generate appropriate answers to new questions. Step 3: The generation unit outputs the answers generated by the analysis unit as video and audio. The generation unit can generate video using deepfake technology and audio using speech synthesis technology. Furthermore, the generation unit can generate video and audio in real time. For example, it can generate video of a specific person using deepfake technology and generate audio by mimicking the voice of a specific person using speech synthesis technology. Real-time generation allows for the instantaneous generation of video and audio in response to user questions.

[0096] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0097] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0098] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0099] Each of the multiple elements described above, including the collection unit, analysis unit, and generation unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the collection unit collects data on a specific person using the camera 42 and microphone 38B of the smart device 14 and uses the control unit 46A to train the AI. The analysis unit is implemented in the identification processing unit 290 of the data processing unit 12, which analyzes the user's questions based on the collected data and generates appropriate answers. The generation unit is implemented in the control unit 46A of the smart device 14, which generates video and audio of a specific person using deepfake technology or speech synthesis technology. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0100] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0101] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0102] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0103] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0104] The microphone 238 receives voice commands and other instructions from the user by receiving voice signals. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0105] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0106] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0107] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0108] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0109] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0110] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0111] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0112] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0113] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0114] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0115] Each of the multiple elements described above, including the collection unit, analysis unit, and generation unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the collection unit collects data on a specific person using the camera 42 and microphone 238 of the smart glasses 214 and uses the control unit 46A to train the AI. The analysis unit is implemented in the identification processing unit 290 of the data processing unit 12, which analyzes the user's questions based on the collected data and generates appropriate answers. The generation unit is implemented in the control unit 46A of the smart glasses 214, which generates images and sounds of a specific person using deepfake technology and speech synthesis technology. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0116] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0117] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0118] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0119] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0120] The microphone 238 receives voice commands and other instructions from the user by receiving voice signals. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0121] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0122] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0123] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0124] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0125] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0126] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0127] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0128] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0129] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0130] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0131] Each of the multiple elements described above, including the collection unit, analysis unit, and generation unit, is implemented in at least one of the following: the headset terminal 314 and the data processing unit 12. For example, the collection unit collects data on a specific person using the camera 42 and microphone 238 of the headset terminal 314 and uses the control unit 46A to train the AI. The analysis unit is implemented in the identification processing unit 290 of the data processing unit 12, which analyzes the user's questions based on the collected data and generates appropriate answers. The generation unit is implemented in the control unit 46A of the headset terminal 314, which generates video and audio of a specific person using deepfake technology and speech synthesis technology. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0132] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0133] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0134] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0135] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0136] The microphone 238 receives voice commands and other instructions from the user by receiving voice signals. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0137] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0138] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0139] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0140] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0141] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0142] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0143] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0144] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0145] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0146] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0147] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0148] Each of the multiple elements described above, including the collection unit, analysis unit, and generation unit, is implemented in, for example, at least one of the robot 414 and the data processing unit 12. For example, the collection unit collects data on a specific person using the camera 42 and microphone 238 of the robot 414 and uses the control unit 46A to train the AI. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the user's questions based on the collected data and generates appropriate answers. The generation unit is implemented, for example, by the control unit 46A of the robot 414, which generates images and sounds of a specific person using deepfake technology or speech synthesis technology. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0149] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0150] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0151] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0152] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0153] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0154] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0155] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0156] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0157] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0158] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0159] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0160] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0161] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0162] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0163] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0164] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0165] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0166] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0167] (Note 1) A data collection unit that collects data on specific individuals and uses it to train an AI, An analysis unit analyzes user questions and inquiries based on the data collected by the aforementioned collection unit and generates answers. The system includes a generation unit that outputs the response generated by the analysis unit as video and audio. A system characterized by the following features. (Note 2) The system according to Appendix 1, characterized in that the collection unit collects data from interview videos or recordings of lectures of specific individuals. (Note 3) The aforementioned analysis unit, Analyze user questions using natural language processing techniques. The system described in Appendix 1, characterized by the features described herein. (Note 4) The system according to Appendix 1, characterized in that the generation unit generates video or audio in real time using deepfake technology or speech synthesis technology. (Note 5) The aforementioned collection unit is It estimates the user's emotions and adjusts the types of data collected based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned collection unit is Analyze the background and context of a specific person's statements to determine the priority of data to collect. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned collection unit is Adjust the level of detail in the data collected based on the frequency and importance of statements made by specific individuals. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned collection unit is It estimates the user's emotions and adjusts the timing of data collection based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned collection unit is Select the data to collect by considering the geographical context of the statements made by specific individuals. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned collection unit is We will also collect and compare statements from other prominent figures related to the statements of a specific individual. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned analysis unit, It estimates the user's emotions and adjusts the analysis algorithm based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned analysis unit, Analyze the background information of the question and apply analytical methods to generate more specific answers. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, Applying different analysis algorithms depending on the question category improves the accuracy of the answers. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, It estimates the user's emotions and adjusts how the analysis results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, We will prioritize the analysis based on when the questions were submitted. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, Adjust the order of analysis based on the relevance of the questions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The generating unit is It estimates the user's emotions and adjusts the way it generates video and audio based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 18) The generating unit is Analyze the tone and pace of a specific person's speech and reflect it in the generated video and audio. The system described in Appendix 1, characterized by the features described herein. (Note 19) The generating unit is Adjust the level of detail in the generated video and audio based on the context of a specific person's statements. The system described in Appendix 1, characterized by the features described herein. (Note 20) The generating unit is It estimates the user's emotions and adjusts the length of the generated video and audio based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 21) The generating unit is Customize the generated video and audio by taking into account the geographical context of a specific person's statements. The system described in Appendix 1, characterized by the features described herein. (Note 22) The generating unit is The system also generates and displays statements from other prominent figures related to a specific person's statements for comparison. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]

[0168] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A device comprising a processor and a memory for storing a program executed by the processor, The processor executes the program, A collection unit collects data on a specific person, including at least one of the following: interview videos of the specific person, recordings of the specific person's lectures, and the contents of books by the specific person; and uses the data on the specific person to perform a learning process on an AI, which is a machine learning model, so that the AI ​​imitates at least one of the patterns of the specific person's speech and actions. An analysis unit, which takes at least one of the user's questions and consultations as input to the AI ​​learned by the collection unit, causes the AI ​​to analyze at least one of the questions and consultations and generate an answer. A generation unit that outputs the response generated by the analysis unit as video or audio, and functions as such, The aforementioned collection unit is Analyze the background and context of a specific person's statements to determine the priority of data to collect. A system characterized by the following features.

2. The aforementioned analysis unit, Analyze user questions using natural language processing techniques. The system according to feature 1.

3. The aforementioned collection unit is The system estimates the user's emotions and, if it estimates the user is stressed, adjusts the types of data collected based on that estimate by prioritizing the collection of interview videos or lecture recordings of specific individuals who can help them relax. The system according to feature 1.