system

The system addresses natural communication challenges in video calls by using eye-tracking correction, simultaneous translation, and lip-syncing technologies to ensure users appear to look at the camera and speak in their own voice, enhancing communication quality.

JP2026072677APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing video call systems face challenges in achieving natural communication due to barriers such as eye line and language differences.

Method used

A system combining eye-tracking correction, simultaneous translation, and lip-syncing technologies to enable users to look at the camera while speaking in their own voice and synchronized lip movements, overcoming language barriers.

Benefits of technology

Enables natural communication in video calls by correcting eye movements, translating speech in real-time, and synchronizing lip movements, eliminating language barriers and improving familiarity and trust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072677000001_ABST
    Figure 2026072677000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to achieve natural communication in video calls. [Solution] The system according to the embodiment comprises an acquisition unit, a correction unit, a voice acquisition unit, a translation unit, and a lip correction unit. The acquisition unit acquires the user's eye movements. The correction unit corrects the eye movements acquired by the acquisition unit to make the user look at the camera. The voice acquisition unit acquires the voice. The translation unit translates the voice acquired by the voice acquisition unit and reads the translated voice aloud in the user's own voice. The lip correction unit corrects the lip movements to match the translated voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, it is difficult to achieve natural communication in a video call, and particularly, the barriers of eye line and language are problems.

[0005] The system according to the embodiment aims to achieve natural communication in a video call.

Means for Solving the Problems

[0006] The system according to the embodiment comprises an acquisition unit, a correction unit, a voice acquisition unit, a translation unit, and a lip correction unit. The acquisition unit acquires the user's eye movements. The correction unit corrects the eye movements acquired by the acquisition unit to make the user look directly at the camera. The voice acquisition unit acquires the voice. The translation unit translates the voice acquired by the voice acquisition unit and reads the translated voice aloud in the user's own voice. The lip correction unit corrects the lip movements to match the translated voice. [Effects of the Invention]

[0007] The system according to this embodiment can enable natural communication in video calls. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The next-generation video call system according to an embodiment of the present invention is a system that realizes natural communication by combining eye-tracking correction, simultaneous translation, voice personalization, and lip-syncing technologies. The next-generation video call system allows users to speak while looking the other party in the eye without worrying about language barriers, solving challenges in business meetings, online education, and international exchange. With the spread of remote work and the advancement of AI technology, now is a prime time for adoption, and the market size is projected to exceed US$94.07 billion by 2032. The next-generation video call system is composed of the following technologies combined. For example, eye-tracking correction technology tracks the user's eye movements in real time and corrects them to look at the camera. This allows the user to converse while looking the other party in the eye, improving familiarity and trust. Simultaneous translation + text-to-speech + voice personalization technology translates speech in real time and reads the translated words in the user's own voice. This eliminates language barriers and improves convenience in international business and meetings. For example, when a user who only speaks Japanese converses with someone who speaks English, the system translates Japanese into English and reads the English speech in a Japanese voice. Lip-sync technology modifies lip movements to match the translated language. This ensures that lip movements and speech are synchronized, eliminating visual inconsistencies. For example, when translating the Japanese "konnichiwa" to the English "Hello," the lip movements are also modified to match "Hello." By combining these technologies, next-generation video conferencing systems enable users to communicate naturally without language barriers. It is expected to be used in a variety of scenarios, including business meetings, online education, and international exchange. This allows next-generation video conferencing systems to achieve natural communication without users feeling the language barrier.

[0029] The next-generation video call system according to this embodiment comprises an acquisition unit, a correction unit, an audio acquisition unit, a translation unit, and a lip correction unit. The acquisition unit acquires the user's eye movements. The acquisition unit can, for example, track the user's eye movements in real time using a camera. The acquisition unit can also detect eye movements using an infrared sensor. Furthermore, the acquisition unit can use a combination of multiple sensors to acquire the user's eye movements with high accuracy. For example, the acquisition unit can track the user's eye movements with high accuracy by combining a camera and an infrared sensor. The correction unit corrects the eye movements acquired by the acquisition unit to match the camera's gaze. The correction unit uses, for example, an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze. Furthermore, the correction unit can also correct the user's eye movements in real time. Furthermore, the correction unit can adjust the timing of the correction to make the user's eye movements appear natural. For example, the correction unit uses an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze. The audio acquisition unit acquires the user's voice. The audio acquisition unit can, for example, acquire the user's voice in real time using a microphone. Furthermore, the voice acquisition unit can acquire clear audio using noise cancellation technology. In addition, the voice acquisition unit can use multiple microphones in combination to acquire the user's voice with high accuracy. For example, the voice acquisition unit can acquire the user's voice with high accuracy by combining microphones and noise cancellation technology. The translation unit translates the audio acquired by the voice acquisition unit and reads the translated audio aloud in the user's voice. The translation unit can, for example, use speech recognition technology to convert the user's voice into text and then translate that text. The translation unit can also use speech synthesis technology to read the translated audio aloud in the user's voice. Furthermore, the translation unit can learn the user's voice and read aloud in a more natural voice. For example, the translation unit can use speech recognition technology to convert the user's voice into text and then translate that text. The lip correction unit corrects the lip movements to match the translated audio. The lip correction unit can, for example, track the user's lip movements in real time and correct them to match the translated audio.Furthermore, the lip correction unit can adjust the timing of corrections to make the user's lip movements appear more natural. In addition, the lip correction unit can use a combination of multiple sensors to correct the user's lip movements with high precision. For example, the lip correction unit can combine a camera and an infrared sensor to track the user's lip movements with high precision and correct them in accordance with the translated audio. As a result, the next-generation video call system according to this embodiment can correct and translate the user's eye movements and voice in real time, enabling natural communication.

[0030] The acquisition unit captures the user's eye movements. For example, the acquisition unit can track the user's eye movements in real time using a camera. Specifically, the camera captures the user's eye movements at high resolution and processes the data in real time. Furthermore, the acquisition unit can also detect eye movements using an infrared sensor. Infrared sensors can accurately detect eye movements even in dark environments, and when combined with a camera, they enable more accurate tracking. The acquisition unit can also use multiple sensors in combination to acquire the user's eye movements with high accuracy. For example, combining a camera and an infrared sensor allows for highly accurate tracking of the user's eye movements. This enables the acquisition unit to accurately understand the user's eye movements and provide the data necessary for subsequent processing. Furthermore, the acquisition unit can analyze the user's eye movements in real time to identify the direction of gaze and the point of fixation. This allows for accurate understanding of where the user is looking and can be reflected in subsequent processing. For example, the acquisition unit can identify which part of the screen the user is looking at and use that information to correct camera gaze or perform other functions. Furthermore, the acquisition unit can learn the user's eye movement patterns and automatically adjust the tracking settings to be optimal for each individual user. This allows the acquisition unit to acquire the user's eye movements with high accuracy and efficiency, improving the overall system performance.

[0031] The correction unit corrects the eye movements acquired by the acquisition unit to match the camera's gaze. For example, the correction unit uses an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze. Specifically, the correction unit calculates the user's gaze angle in real time and corrects it to match the camera's gaze based on that information. The correction unit can also correct the user's eye movements in real time. This makes it appear to the other party that the user is looking directly at the camera when they are looking at the screen. Furthermore, the correction unit can adjust the timing of the correction to make the user's eye movements appear more natural. For example, the correction unit uses an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze. This corrects the user's eye movements to appear natural and does not cause discomfort to the other party. The correction unit can also learn the user's eye movement patterns and automatically adjust the optimal correction settings for each individual user. This allows the correction unit to correct the user's eye movements with high accuracy and naturalness, improving the overall system performance. Furthermore, the correction unit can consider not only the user's eye movements but also the movements and expressions of the entire face when performing corrections. This adjusts the user's facial expressions and gaze to appear more natural, enabling more realistic communication. For example, the correction unit analyzes facial movements and expressions and corrects the gaze based on that analysis to make the user appear more natural. As a result, the correction unit can correct the user's eye movements with high precision and in a natural way, improving the overall performance of the system.

[0032] The voice acquisition unit acquires the user's voice. For example, the voice acquisition unit can acquire the user's voice in real time using a microphone. Specifically, the voice acquisition unit uses a high-sensitivity microphone to clearly capture the user's voice. The voice acquisition unit can also acquire clear audio using noise cancellation technology. Noise cancellation technology removes ambient noise, allowing for high-precision acquisition of only the user's voice. Furthermore, the voice acquisition unit can use multiple microphones in combination to acquire the user's voice with high precision. For example, the voice acquisition unit can combine microphones and noise cancellation technology to acquire the user's voice with high precision. This allows the voice acquisition unit to acquire the user's voice clearly and with high precision, providing the data necessary for subsequent processing. In addition, the voice acquisition unit can analyze the characteristics of the user's voice and make adjustments to improve the audio quality. For example, the voice acquisition unit analyzes the tone and pitch of the user's voice and optimizes the audio quality based on that analysis. The voice acquisition unit can also learn the user's voice patterns and automatically adjust the optimal voice acquisition settings for each individual user. This allows the voice acquisition unit to acquire the user's voice with high accuracy and efficiency, improving the overall system performance. Furthermore, the voice acquisition unit can analyze the user's voice in real time and provide feedback to improve the quality of the voice. This enables the voice acquisition unit to acquire the user's voice with high accuracy and efficiency, improving the overall system performance.

[0033] The translation unit translates the audio acquired by the speech acquisition unit and reads the translated audio aloud in the user's own voice. For example, the translation unit can convert the user's voice into text using speech recognition technology and then translate that text. Specifically, the translation unit uses speech recognition technology to convert the user's voice into text with high accuracy and then translates that text into multiple languages. Furthermore, the translation unit can also use speech synthesis technology to read the translated audio aloud in the user's own voice. Speech synthesis technology learns the user's voice characteristics and can read the translated audio in a natural voice. Moreover, the translation unit can learn the user's voice characteristics and read in an even more natural voice. For example, the translation unit uses speech recognition technology to convert the user's voice into text and then translates that text. This allows the translation unit to translate the user's voice with high accuracy and read it aloud in a natural voice. Furthermore, the translation unit can analyze the characteristics of the user's voice and make adjustments to improve the audio quality. For example, the translation unit analyzes the tone and pitch of the user's voice and optimizes the audio quality based on that analysis. Furthermore, the translation unit can learn the user's voice patterns and automatically adjust the optimal speech synthesis settings for each individual user. This allows the translation unit to translate the user's voice with high accuracy and efficiency, improving the overall system performance. In addition, the translation unit can analyze the user's voice in real time and provide feedback to improve the voice quality. This allows the translation unit to translate the user's voice with high accuracy and efficiency, improving the overall system performance.

[0034] The lip correction unit corrects lip movements to match translated speech. For example, it tracks the user's lip movements in real time and corrects them to match the translated speech. Specifically, the lip correction unit captures the user's lip movements using a high-resolution camera and processes the data in real time. The lip correction unit can also adjust the timing of corrections to make the user's lip movements appear natural. For example, it tracks the user's lip movements in real time and corrects them to match the translated speech. This corrects the user's lip movements to look natural and avoid causing discomfort to the listener. Furthermore, the lip correction unit can use a combination of sensors to correct the user's lip movements with high precision. For example, it can combine a camera and an infrared sensor to track the user's lip movements with high precision and correct them to match the translated speech. This allows the lip correction unit to correct the user's lip movements with high precision and naturalness, improving the overall system performance. Additionally, the lip correction unit can learn the user's lip movement patterns and automatically adjust the optimal correction settings for each individual user. This allows the lip correction unit to correct the user's lip movements with high precision and efficiency, improving the overall system performance. Furthermore, the lip correction unit can analyze the user's lip movements in real time and provide feedback to improve the quality of the correction. This enables the lip correction unit to correct the user's lip movements with high precision and efficiency, improving the overall system performance.

[0035] The acquisition unit can track the user's eye movements in real time. For example, the acquisition unit can track the user's eye movements in real time using a camera. The acquisition unit can also detect eye movements using an infrared sensor. Furthermore, the acquisition unit can use multiple sensors in combination to acquire the user's eye movements with high accuracy. For example, the acquisition unit can combine a camera and an infrared sensor to track the user's eye movements with high accuracy. This enables more accurate gaze correction by tracking the user's eye movements in real time. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's eye movement data acquired by the camera into a generating AI and have the generating AI perform the processing of tracking in real time.

[0036] The correction unit can correct the eye movements acquired by the acquisition unit to match the camera's gaze. For example, the correction unit uses an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze. The correction unit can also correct the user's eye movements in real time. Furthermore, the correction unit can adjust the timing of the correction to make the user's eye movements appear more natural. For example, the correction unit uses an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze. This allows the user to look the other person in the eye while talking by correcting the eye movements to match the camera's gaze. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input the eye movement data acquired by the acquisition unit into a generating AI and have the generating AI perform the process of correcting it to match the camera's gaze.

[0037] The voice acquisition unit can acquire the user's voice in real time. For example, the voice acquisition unit can acquire the user's voice in real time using a microphone. The voice acquisition unit can also acquire clear audio using noise cancellation technology. Furthermore, the voice acquisition unit can use multiple microphones in combination to acquire the user's voice with high accuracy. For example, the voice acquisition unit can acquire the user's voice with high accuracy by combining a microphone and noise cancellation technology. This enables instant translation and reading aloud by acquiring the user's voice in real time. Some or all of the above processing in the voice acquisition unit may be performed using AI, for example, or without AI. For example, the voice acquisition unit can input the user's voice data acquired by the microphone into a generating AI and have the generating AI execute the process of acquiring the voice in real time.

[0038] The translation unit can translate the audio acquired by the audio acquisition unit and read the translated audio aloud in the user's own voice. For example, the translation unit can use speech recognition technology to convert the user's voice into text and then translate that text. The translation unit can also use speech synthesis technology to read the translated audio aloud in the user's own voice. Furthermore, the translation unit can learn the user's voice and read it aloud in a more natural voice. For example, the translation unit can use speech recognition technology to convert the user's voice into text and then translate that text. By translating the audio and reading it aloud in the user's own voice, language barriers are eliminated, and natural communication is achieved. Some or all of the above processing in the translation unit may be performed using AI, for example, or without AI. For example, the translation unit can input the audio data acquired by the audio acquisition unit into a generating AI and have the generating AI perform the translation and speech synthesis processes.

[0039] The lip correction unit can correct lip movements to match translated speech. For example, the lip correction unit can track the user's lip movements in real time and correct them to match the translated speech. The lip correction unit can also adjust the timing of the corrections to make the user's lip movements appear more natural. Furthermore, the lip correction unit can use a combination of multiple sensors to correct the user's lip movements with high precision. For example, the lip correction unit can combine a camera and an infrared sensor to track the user's lip movements with high precision and correct them to match the translated speech. This eliminates visual inconsistencies by correcting lip movements to match the translated speech. Some or all of the above processing in the lip correction unit may be performed using AI, for example, or without AI. For example, the lip correction unit can input data on the user's lip movements into a generating AI and have the generating AI perform the process of correcting the lip movements in real time.

[0040] The acquisition unit can analyze the user's past eye movement data and select the optimal acquisition method. For example, the acquisition unit can identify points that the user has frequently looked at in the past and prioritize acquiring eye movements to those points. The acquisition unit can also analyze the user's past eye movement data to identify eye movement patterns in specific time periods and adjust the acquisition method. Furthermore, the acquisition unit can predict eye movements under specific circumstances based on the user's past eye movement data and optimize the acquisition method. For example, the acquisition unit can identify points that the user has frequently looked at in the past and prioritize acquiring eye movements to those points. By analyzing past eye movement data, the acquisition unit can select the optimal acquisition method and improve accuracy. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's past eye movement data into a generating AI and have the generating AI perform the process of selecting the optimal acquisition method.

[0041] The acquisition unit can identify the user's current gaze focus when acquiring eye movements and prioritize acquiring gaze towards specific objects. For example, if the user is looking at a specific icon on the screen, the acquisition unit will prioritize acquiring the gaze movement toward that icon. The acquisition unit can also prioritize acquiring the gaze movement toward the face of the person the user is talking to. Furthermore, if the user is reading a document, the acquisition unit can also prioritize acquiring the gaze movement toward that document. For example, if the acquisition unit is looking at a specific icon on the screen, the acquisition unit will prioritize acquiring the gaze movement toward that icon. By identifying the user's gaze focus and prioritizing the acquisition of gaze toward specific objects, more accurate data can be collected. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's gaze focus data into a generating AI and cause the generating AI to execute the process of prioritizing the acquisition of gaze toward specific objects.

[0042] The acquisition unit can adjust the acquisition accuracy when acquiring eye movements, taking into account changes in the user's ambient light. For example, if the ambient light is bright, the acquisition unit can increase the acquisition accuracy to acquire detailed eye movements. Conversely, if the ambient light is dim, the acquisition unit can also decrease the acquisition accuracy to reduce noise. Furthermore, if the ambient light fluctuates, the acquisition unit can adjust the acquisition accuracy in real time. For example, if the ambient light is bright, the acquisition unit can increase the acquisition accuracy to acquire detailed eye movements. By adjusting the acquisition accuracy to take into account changes in ambient light, more accurate eye movement data can be acquired. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input ambient light data into a generating AI and have the generating AI perform the process of adjusting the acquisition accuracy.

[0043] The acquisition unit can correct the direction of the gaze in conjunction with the user's head movements when acquiring eye movements. For example, the acquisition unit corrects the direction of the gaze in real time when the user moves their head. The acquisition unit can also correct the angle of the gaze to acquire accurate data when the user tilts their head. Furthermore, the acquisition unit can correct the height of the gaze when the user moves their head up and down. For example, the acquisition unit corrects the direction of the gaze in real time when the user moves their head. This allows for the acquisition of more accurate gaze data by correcting the direction of the gaze in conjunction with the user's head movements. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input data of the user's head movements into a generating AI and have the generating AI perform the process of correcting the direction of the gaze.

[0044] The correction unit analyzes the user's facial expression during correction and can correct the gaze while maintaining a natural expression. For example, when the user is smiling, the correction unit can maintain the smile while correcting the gaze. The correction unit can also maintain a serious expression when the user is serious while correcting the gaze. Furthermore, the correction unit can maintain a surprised expression when the user is surprised while correcting the gaze. For example, when the user is smiling, the correction unit can maintain the smile while correcting the gaze. In this way, by analyzing the user's facial expression and correcting the gaze while maintaining a natural expression, visual unnaturalness is reduced. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input the user's facial expression data into a generating AI and cause the generating AI to perform the process of correcting the gaze while maintaining a natural expression.

[0045] The correction unit can optimize the timing of correction by taking into account the speed of the user's eye movements. For example, if the user's eye movements are fast, the correction unit can shorten the correction timing to improve real-time performance. Conversely, if the user's eye movements are slow, the correction unit can extend the correction timing to maintain natural movement. Furthermore, if the user's eye movements are irregular, the correction unit can dynamically adjust the correction timing. For example, if the user's eye movements are fast, the correction unit can shorten the correction timing to improve real-time performance. This optimizes the correction timing by taking into account the speed of the user's eye movements, thereby achieving more natural gaze correction. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input user eye movement data into a generating AI and have the generating AI perform the process of optimizing the correction timing.

[0046] The correction unit can adjust the correction accuracy during correction, taking into account whether the user is wearing glasses or contact lenses. For example, if the user is wearing glasses, the correction unit can increase the correction accuracy to maintain the accuracy of the gaze. The correction unit can also adjust the correction accuracy to maintain the naturalness of the gaze if the user is wearing contact lenses. Furthermore, if the user is not wearing glasses or contact lenses, the correction unit can optimize the correction accuracy to achieve both accuracy and naturalness of the gaze. For example, if the user is wearing glasses, the correction unit can increase the correction accuracy to maintain the accuracy of the gaze. This maintains the accuracy of the gaze by adjusting the correction accuracy considering whether the user is wearing glasses or contact lenses. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input data on whether the user is wearing glasses or contact lenses into a generating AI and have the generating AI perform the process of adjusting the correction accuracy.

[0047] The correction unit can correct the user's gaze in conjunction with the orientation of their face during correction. For example, when the user moves their face from side to side, the correction unit maintains the orientation of the face while correcting the gaze. The correction unit can also maintain the orientation of the face while correcting the gaze when the user moves their face up and down. Furthermore, the correction unit can maintain the orientation of the face while correcting the gaze when the user moves their face diagonally. For example, when the user moves their face from side to side, the correction unit maintains the orientation of the face while correcting the gaze. This results in a more natural gaze by correcting the gaze in conjunction with the orientation of the user's face. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input data on the orientation of the user's face into a generating AI and have the generating AI execute the process of correcting the gaze.

[0048] The correction unit can correct the user's gaze in conjunction with the orientation of their face during correction. For example, when the user moves their face from side to side, the correction unit maintains the orientation of the face while correcting the gaze. The correction unit can also maintain the orientation of the face while correcting the gaze when the user moves their face up and down. Furthermore, the correction unit can maintain the orientation of the face while correcting the gaze when the user moves their face diagonally. For example, when the user moves their face from side to side, the correction unit maintains the orientation of the face while correcting the gaze. This results in a more natural gaze by correcting the gaze in conjunction with the orientation of the user's face. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input data on the orientation of the user's face into a generating AI and have the generating AI execute the process of correcting the gaze.

[0049] The voice acquisition unit can analyze the user's past voice data and select the optimal acquisition method. For example, the voice acquisition unit can identify phrases that the user has frequently uttered in the past and prioritize acquiring audio for those phrases. The voice acquisition unit can also analyze the user's past voice data to identify patterns in speech during specific time periods and adjust the acquisition method accordingly. Furthermore, the voice acquisition unit can predict speech under specific circumstances based on the user's past voice data and optimize the acquisition method. For example, the voice acquisition unit can identify phrases that the user has frequently uttered in the past and prioritize acquiring audio for those phrases. By analyzing past voice data, the optimal acquisition method is selected, improving accuracy. Some or all of the above processing in the voice acquisition unit may be performed using AI, for example, or without AI. For example, the voice acquisition unit can input the user's past voice data into a generating AI and have the generating AI perform the process of selecting the optimal acquisition method.

[0050] The voice acquisition unit can acquire clear audio by filtering out ambient noise around the user during voice acquisition. For example, if the user is in a noisy environment, the voice acquisition unit can acquire clear audio using noise cancellation technology. Furthermore, if the user is in a quiet environment, the voice acquisition unit can acquire natural audio with minimal noise filtering. Additionally, if the user is moving, the voice acquisition unit can acquire clear audio by filtering out ambient noise in real time. For example, if the user is in a noisy environment, the voice acquisition unit can acquire clear audio using noise cancellation technology. This filters out ambient noise, resulting in clear audio. Some or all of the above processing in the voice acquisition unit may be performed using AI, or without AI. For example, the voice acquisition unit can input ambient noise data into a generating AI and have the generating AI perform the noise filtering process.

[0051] The audio acquisition unit can adjust the acquisition accuracy when acquiring audio, taking into account the user's microphone position. For example, if the user places the microphone close by, the audio acquisition unit can increase the acquisition accuracy to acquire detailed audio. Conversely, if the user places the microphone far away, the audio acquisition unit can also adjust the acquisition accuracy to acquire clear audio. Furthermore, if the user moves the microphone, the audio acquisition unit can adjust the acquisition accuracy in real time. For example, if the user places the microphone close by, the audio acquisition unit can increase the acquisition accuracy to acquire detailed audio. This ensures clear audio is acquired by adjusting the acquisition accuracy considering the microphone position. Some or all of the above processing in the audio acquisition unit may be performed using AI, for example, or without AI. For example, the audio acquisition unit can input microphone position data into a generating AI and have the generating AI perform the process of adjusting the acquisition accuracy.

[0052] The voice acquisition unit can optimize its voice acquisition method according to the user's speaking speed. For example, if the user speaks quickly, the voice acquisition unit adjusts the acquisition method to improve real-time performance. The voice acquisition unit can also adjust the acquisition method to acquire more detailed audio if the user speaks slowly. Furthermore, if the user speaks at an irregular speed, the voice acquisition unit can dynamically adjust the acquisition method to acquire the optimal audio. For example, if the user speaks quickly, the voice acquisition unit adjusts the acquisition method to improve real-time performance. This optimizes the acquisition method according to the user's speaking speed, resulting in the acquisition of more accurate audio data. Some or all of the above-described processes in the voice acquisition unit may be performed using AI, for example, or without AI. For example, the voice acquisition unit can input user speaking speed data into a generating AI and have the generating AI perform the process of optimizing the acquisition method.

[0053] The translation unit can improve translation accuracy by taking into account the user's specialized terminology and slang during translation. For example, if the user frequently uses specialized terminology, the translation unit will accurately translate that terminology. The translation unit can also appropriately translate slang if the user uses it. Furthermore, if the user uses specific industry-specific terminology, the translation unit can accurately translate that industry-specific terminology. For example, if the user frequently uses specialized terminology, the translation unit will accurately translate that terminology. This improves translation accuracy by considering specialized terminology and slang. Some or all of the above processing in the translation unit may be performed using AI, for example, or without AI. For example, the translation unit can input data on the user's specialized terminology and slang into a generating AI and have the generating AI perform processing to improve translation accuracy.

[0054] The translation unit can optimize the timing of translation according to the user's speaking speed. For example, if the user speaks quickly, the translation unit can shorten the translation timing to improve real-time performance. Conversely, if the user speaks slowly, the translation unit can extend the translation timing to provide a more detailed translation. Furthermore, the translation unit can dynamically adjust the translation timing if the user speaks at an irregular pace. For example, if the user speaks quickly, the translation unit can shorten the translation timing to improve real-time performance. This improves real-time performance by optimizing the translation timing according to the user's speaking speed. Some or all of the above processing in the translation unit may be performed using AI, for example, or without AI. For example, the translation unit can input user speaking speed data into a generating AI and have the generating AI perform the process of optimizing the translation timing.

[0055] The translation unit can adjust translation accuracy by taking into account regional differences in the user's spoken language during translation. For example, if the user speaks a dialect from a specific region, the translation unit will translate that dialect appropriately. The translation unit can also take into account regional differences if the user speaks languages ​​from different regions. Furthermore, if the user speaks a mix of languages ​​from multiple regions, the translation unit can also take into account regional differences. For example, if the translation unit speaks a dialect from a specific region, the translation unit will translate that dialect appropriately. This provides a more accurate translation by taking regional differences into account. Some or all of the above processing in the translation unit may be performed using AI, for example, or not using AI. For example, the translation unit can input regional difference data of the user's spoken language into a generating AI and have the generating AI perform the process of adjusting translation accuracy.

[0056] The translation unit can apply different translation algorithms depending on the category of what the user is saying during translation. For example, if the user uses business terminology, the translation unit will apply a business translation algorithm. It can also apply a conversational translation algorithm if the user is engaging in everyday conversation. Furthermore, if the user uses technical terminology, the translation unit can apply a technical translation algorithm. This improves translation accuracy by applying translation algorithms appropriate to the content category. Some or all of the above processing in the translation unit may be performed using AI, for example, or without AI. For example, the translation unit can input category data of what the user is saying into a generating AI and have the generating AI perform the process of applying different translation algorithms.

[0057] The lip correction unit analyzes the user's facial expression during correction and can correct lip movements while maintaining a natural expression. For example, when the user is smiling, the lip correction unit can maintain the smile by correcting lip movements. It can also maintain a serious expression when the user is serious by correcting lip movements. Furthermore, it can maintain a surprised expression when the user is surprised by correcting lip movements. For example, when the user is smiling, the lip correction unit can maintain the smile by correcting lip movements. This reduces visual unnaturalness by analyzing the user's facial expression and correcting lip movements while maintaining a natural expression. Some or all of the above processing in the lip correction unit may be performed using AI, for example, or without AI. For example, the lip correction unit can input the user's facial expression data into a generating AI and cause the generating AI to perform the process of correcting lip movements while maintaining a natural expression.

[0058] The lip correction unit can optimize the timing of lip movement correction according to the user's speaking speed during correction. For example, if the user speaks quickly, the lip correction unit can shorten the timing of lip movement correction to improve real-time performance. Conversely, if the user speaks slowly, the lip correction unit can extend the timing of lip movement correction to maintain natural movement. Furthermore, if the user speaks at an irregular speed, the lip correction unit can dynamically adjust the timing of lip movement correction. For example, if the user speaks quickly, the lip correction unit can shorten the timing of lip movement correction to improve real-time performance. This optimizes the timing of lip movement correction according to the user's speaking speed, resulting in more natural movement. Some or all of the above processing in the lip correction unit may be performed using AI, for example, or without AI. For example, the lip correction unit can input user speaking speed data into a generating AI and have the generating AI perform the process of optimizing the timing of lip movement correction.

[0059] The lip correction unit can correct lip movement in conjunction with the user's facial orientation during correction. For example, when the user moves their face from side to side, the lip correction unit can correct lip movement while maintaining the facial orientation. It can also correct lip movement while maintaining the facial orientation when the user moves their face up and down. Furthermore, it can correct lip movement while maintaining the facial orientation when the user moves their face diagonally. For example, when the user moves their face from side to side, the lip correction unit corrects lip movement while maintaining the facial orientation. This allows for more natural movement by correcting lip movement in conjunction with the user's facial orientation. Some or all of the above-described processes in the lip correction unit may be performed using AI, or without AI. For example, the lip correction unit can input user facial orientation data into a generating AI and have the generating AI perform the process of correcting lip movement.

[0060] The lip correction unit can apply different lip movement correction algorithms depending on the category of what the user is saying during correction. For example, if the user is using business terminology, the lip correction unit can apply a business-oriented lip movement correction algorithm. It can also apply a lip movement correction algorithm for everyday conversation if the user is engaging in everyday conversation. Furthermore, if the user is using technical terminology, the lip correction unit can apply a technical-oriented lip movement correction algorithm. For example, if the user is using business terminology, the lip correction unit can apply a business-oriented lip movement correction algorithm. This allows for more natural movement by applying a lip movement correction algorithm appropriate to the content category. Some or all of the above processing in the lip correction unit may be performed using AI, or without AI. For example, the lip correction unit can input category data of what the user is saying into a generating AI and have the generating AI perform the process of applying different lip movement correction algorithms.

[0061] The lip correction unit can optimize the method of correcting lip movements according to the user's speaking speed during correction. For example, if the user speaks quickly, the lip correction unit adjusts the method of correcting lip movements to improve real-time performance. The lip correction unit can also adjust the method of correcting lip movements to maintain natural movement if the user speaks slowly. Furthermore, the lip correction unit can dynamically adjust the method of correcting lip movements if the user speaks at an irregular speed. For example, if the user speaks quickly, the lip correction unit adjusts the method of correcting lip movements to improve real-time performance. This optimizes the method of correcting lip movements according to the user's speaking speed, resulting in more natural movement. Some or all of the above processing in the lip correction unit may be performed using AI, for example, or without AI. For example, the lip correction unit can input user speaking speed data into a generating AI and have the generating AI perform the process of optimizing the method of correcting lip movements.

[0062] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0063] Next-generation video conferencing systems can also include a background recognition unit that automatically recognizes the user's background and switches to an appropriate virtual background. For example, if the user is at home, the background recognition unit can switch to a virtual office background. It can also switch to a virtual cafe background if the user is out and about. Furthermore, if the user is in a meeting, the background recognition unit can switch to a virtual conference room background. This allows for a more professional impression by automatically recognizing the user's background and switching to an appropriate virtual background.

[0064] Next-generation video conferencing systems can also include a gesture recognition unit that recognizes user gestures and performs actions corresponding to those gestures. For example, if a user raises their hand, the gesture recognition unit can notify the system of a request to speak. It can also highlight a specific object if a user points. Furthermore, if a user claps, the gesture recognition unit can display positive feedback. This allows for more interactive communication by recognizing user gestures and performing actions accordingly.

[0065] Next-generation video conferencing systems can also include an information provision unit that analyzes user speech in real time and provides relevant information. For example, if a user is discussing a specific topic, the information provision unit can display news articles related to that topic. It can also provide answers to questions asked by the user. Furthermore, if a user is referring to specific data, the information provision unit can display details of that data. This allows for improved conversation quality by analyzing user speech in real time and providing relevant information.

[0066] Next-generation video conferencing systems can also include an intent estimation unit that analyzes user speech and estimates the intent behind their statements. For example, if a user asks a question, the intent estimation unit can estimate the intent behind that question and provide an appropriate answer. Furthermore, if a user makes a suggestion, the intent estimation unit can estimate the intent behind that suggestion and provide relevant information. Additionally, if a user expresses an opinion, the intent estimation unit can estimate the intent behind that opinion and provide appropriate feedback. This allows for more effective communication by analyzing user speech and estimating their intent.

[0067] Next-generation video conferencing systems can also include a text display unit that translates user speech in real time and displays the translation as text. For example, if a user is speaking in English, the text display unit will translate their speech into Japanese and display it as text. It can also translate a user's speech into English if they are speaking in Japanese and display it as text. Furthermore, if a user is using multiple languages, the text display unit can translate their speech into each language and display it as text. This allows for real-time translation of user speech and display of the translation as text, eliminating language barriers and facilitating smoother communication.

[0068] The following briefly describes the processing flow for example form 1.

[0069] Step 1: The acquisition unit acquires the user's eye movements. The acquisition unit can, for example, track the user's eye movements in real time using a camera. It can also detect eye movements using an infrared sensor. Furthermore, multiple sensors can be used in combination. For example, a camera and an infrared sensor can be combined to track the user's eye movements with high precision. Step 2: The correction unit corrects the eye movements acquired by the acquisition unit to match the camera's gaze. The correction unit uses an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze, for example. It can also correct the user's eye movements in real time. Furthermore, the timing of the correction can be adjusted. Step 3: The voice acquisition unit acquires the user's voice. The voice acquisition unit can acquire the user's voice in real time, for example, using a microphone. It can also acquire clear audio using noise cancellation technology. Furthermore, multiple microphones can be used in combination. For example, a microphone and noise cancellation technology can be combined to acquire the user's voice with high accuracy. Step 4: The translation unit translates the audio acquired by the audio acquisition unit and reads the translated audio aloud in the user's voice. The translation unit can, for example, use speech recognition technology to convert the user's voice into text and then translate that text. It can also use speech synthesis technology to read the translated audio aloud in the user's voice. Furthermore, it can learn the user's voice and read aloud in a more natural voice. Step 5: The lip correction unit corrects lip movements to match the translated audio. For example, the lip correction unit tracks the user's lip movements in real time and corrects them to match the translated audio. The timing of the correction can also be adjusted. Furthermore, multiple sensors can be used in combination. For example, a camera and an infrared sensor can be combined to track the user's lip movements with high precision and correct them to match the translated audio.

[0070] (Example of form 2) The next-generation video call system according to an embodiment of the present invention is a system that realizes natural communication by combining eye-tracking correction, simultaneous translation, voice personalization, and lip-syncing technologies. The next-generation video call system allows users to speak while looking the other party in the eye without worrying about language barriers, solving challenges in business meetings, online education, and international exchange. With the spread of remote work and the advancement of AI technology, now is a prime time for adoption, and the market size is projected to exceed US$94.07 billion by 2032. The next-generation video call system is composed of the following technologies combined. For example, eye-tracking correction technology tracks the user's eye movements in real time and corrects them to look at the camera. This allows the user to converse while looking the other party in the eye, improving familiarity and trust. Simultaneous translation + text-to-speech + voice personalization technology translates speech in real time and reads the translated words in the user's own voice. This eliminates language barriers and improves convenience in international business and meetings. For example, when a user who only speaks Japanese converses with someone who speaks English, the system translates Japanese into English and reads the English speech in a Japanese voice. Lip-sync technology modifies lip movements to match the translated language. This ensures that lip movements and speech are synchronized, eliminating visual inconsistencies. For example, when translating the Japanese "konnichiwa" to the English "Hello," the lip movements are also modified to match "Hello." By combining these technologies, next-generation video conferencing systems enable users to communicate naturally without language barriers. It is expected to be used in a variety of scenarios, including business meetings, online education, and international exchange. This allows next-generation video conferencing systems to achieve natural communication without users feeling the language barrier.

[0071] The next-generation video call system according to this embodiment comprises an acquisition unit, a correction unit, an audio acquisition unit, a translation unit, and a lip correction unit. The acquisition unit acquires the user's eye movements. The acquisition unit can, for example, track the user's eye movements in real time using a camera. The acquisition unit can also detect eye movements using an infrared sensor. Furthermore, the acquisition unit can use a combination of multiple sensors to acquire the user's eye movements with high accuracy. For example, the acquisition unit can track the user's eye movements with high accuracy by combining a camera and an infrared sensor. The correction unit corrects the eye movements acquired by the acquisition unit to match the camera's gaze. The correction unit uses, for example, an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze. Furthermore, the correction unit can also correct the user's eye movements in real time. Furthermore, the correction unit can adjust the timing of the correction to make the user's eye movements appear natural. For example, the correction unit uses an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze. The audio acquisition unit acquires the user's voice. The audio acquisition unit can, for example, acquire the user's voice in real time using a microphone. Furthermore, the voice acquisition unit can acquire clear audio using noise cancellation technology. In addition, the voice acquisition unit can use multiple microphones in combination to acquire the user's voice with high accuracy. For example, the voice acquisition unit can acquire the user's voice with high accuracy by combining microphones and noise cancellation technology. The translation unit translates the audio acquired by the voice acquisition unit and reads the translated audio aloud in the user's voice. The translation unit can, for example, use speech recognition technology to convert the user's voice into text and then translate that text. The translation unit can also use speech synthesis technology to read the translated audio aloud in the user's voice. Furthermore, the translation unit can learn the user's voice and read aloud in a more natural voice. For example, the translation unit can use speech recognition technology to convert the user's voice into text and then translate that text. The lip correction unit corrects the lip movements to match the translated audio. The lip correction unit can, for example, track the user's lip movements in real time and correct them to match the translated audio.Furthermore, the lip correction unit can adjust the timing of corrections to make the user's lip movements appear more natural. In addition, the lip correction unit can use a combination of multiple sensors to correct the user's lip movements with high precision. For example, the lip correction unit can combine a camera and an infrared sensor to track the user's lip movements with high precision and correct them in accordance with the translated audio. As a result, the next-generation video call system according to this embodiment can correct and translate the user's eye movements and voice in real time, enabling natural communication.

[0072] The acquisition unit captures the user's eye movements. For example, the acquisition unit can track the user's eye movements in real time using a camera. Specifically, the camera captures the user's eye movements at high resolution and processes the data in real time. Furthermore, the acquisition unit can also detect eye movements using an infrared sensor. Infrared sensors can accurately detect eye movements even in dark environments, and when combined with a camera, they enable more accurate tracking. The acquisition unit can also use multiple sensors in combination to acquire the user's eye movements with high accuracy. For example, combining a camera and an infrared sensor allows for highly accurate tracking of the user's eye movements. This enables the acquisition unit to accurately understand the user's eye movements and provide the data necessary for subsequent processing. Furthermore, the acquisition unit can analyze the user's eye movements in real time to identify the direction of gaze and the point of fixation. This allows for accurate understanding of where the user is looking and can be reflected in subsequent processing. For example, the acquisition unit can identify which part of the screen the user is looking at and use that information to correct camera gaze or perform other functions. Furthermore, the acquisition unit can learn the user's eye movement patterns and automatically adjust the tracking settings to be optimal for each individual user. This allows the acquisition unit to acquire the user's eye movements with high accuracy and efficiency, improving the overall system performance.

[0073] The correction unit corrects the eye movements acquired by the acquisition unit to match the camera's gaze. For example, the correction unit uses an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze. Specifically, the correction unit calculates the user's gaze angle in real time and corrects it to match the camera's gaze based on that information. The correction unit can also correct the user's eye movements in real time. This makes it appear to the other party that the user is looking directly at the camera when they are looking at the screen. Furthermore, the correction unit can adjust the timing of the correction to make the user's eye movements appear more natural. For example, the correction unit uses an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze. This corrects the user's eye movements to appear natural and does not cause discomfort to the other party. The correction unit can also learn the user's eye movement patterns and automatically adjust the optimal correction settings for each individual user. This allows the correction unit to correct the user's eye movements with high accuracy and naturalness, improving the overall system performance. Furthermore, the correction unit can consider not only the user's eye movements but also the movements and expressions of the entire face when performing corrections. This adjusts the user's facial expressions and gaze to appear more natural, enabling more realistic communication. For example, the correction unit analyzes facial movements and expressions and corrects the gaze based on that analysis to make the user appear more natural. As a result, the correction unit can correct the user's eye movements with high precision and in a natural way, improving the overall performance of the system.

[0074] The voice acquisition unit acquires the user's voice. For example, the voice acquisition unit can acquire the user's voice in real time using a microphone. Specifically, the voice acquisition unit uses a high-sensitivity microphone to clearly capture the user's voice. The voice acquisition unit can also acquire clear audio using noise cancellation technology. Noise cancellation technology removes ambient noise, allowing for high-precision acquisition of only the user's voice. Furthermore, the voice acquisition unit can use multiple microphones in combination to acquire the user's voice with high precision. For example, the voice acquisition unit can combine microphones and noise cancellation technology to acquire the user's voice with high precision. This allows the voice acquisition unit to acquire the user's voice clearly and with high precision, providing the data necessary for subsequent processing. In addition, the voice acquisition unit can analyze the characteristics of the user's voice and make adjustments to improve the audio quality. For example, the voice acquisition unit analyzes the tone and pitch of the user's voice and optimizes the audio quality based on that analysis. The voice acquisition unit can also learn the user's voice patterns and automatically adjust the optimal voice acquisition settings for each individual user. This allows the voice acquisition unit to acquire the user's voice with high accuracy and efficiency, improving the overall system performance. Furthermore, the voice acquisition unit can analyze the user's voice in real time and provide feedback to improve the quality of the voice. This enables the voice acquisition unit to acquire the user's voice with high accuracy and efficiency, improving the overall system performance.

[0075] The translation unit translates the audio acquired by the speech acquisition unit and reads the translated audio aloud in the user's own voice. For example, the translation unit can convert the user's voice into text using speech recognition technology and then translate that text. Specifically, the translation unit uses speech recognition technology to convert the user's voice into text with high accuracy and then translates that text into multiple languages. Furthermore, the translation unit can also use speech synthesis technology to read the translated audio aloud in the user's own voice. Speech synthesis technology learns the user's voice characteristics and can read the translated audio in a natural voice. Moreover, the translation unit can learn the user's voice characteristics and read in an even more natural voice. For example, the translation unit uses speech recognition technology to convert the user's voice into text and then translates that text. This allows the translation unit to translate the user's voice with high accuracy and read it aloud in a natural voice. Furthermore, the translation unit can analyze the characteristics of the user's voice and make adjustments to improve the audio quality. For example, the translation unit analyzes the tone and pitch of the user's voice and optimizes the audio quality based on that analysis. Furthermore, the translation unit can learn the user's voice patterns and automatically adjust the optimal speech synthesis settings for each individual user. This allows the translation unit to translate the user's voice with high accuracy and efficiency, improving the overall system performance. In addition, the translation unit can analyze the user's voice in real time and provide feedback to improve the voice quality. This allows the translation unit to translate the user's voice with high accuracy and efficiency, improving the overall system performance.

[0076] The lip correction unit corrects lip movements to match translated speech. For example, it tracks the user's lip movements in real time and corrects them to match the translated speech. Specifically, the lip correction unit captures the user's lip movements using a high-resolution camera and processes the data in real time. The lip correction unit can also adjust the timing of corrections to make the user's lip movements appear natural. For example, it tracks the user's lip movements in real time and corrects them to match the translated speech. This corrects the user's lip movements to look natural and avoid causing discomfort to the listener. Furthermore, the lip correction unit can use a combination of sensors to correct the user's lip movements with high precision. For example, it can combine a camera and an infrared sensor to track the user's lip movements with high precision and correct them to match the translated speech. This allows the lip correction unit to correct the user's lip movements with high precision and naturalness, improving the overall system performance. Additionally, the lip correction unit can learn the user's lip movement patterns and automatically adjust the optimal correction settings for each individual user. This allows the lip correction unit to correct the user's lip movements with high precision and efficiency, improving the overall system performance. Furthermore, the lip correction unit can analyze the user's lip movements in real time and provide feedback to improve the quality of the correction. This enables the lip correction unit to correct the user's lip movements with high precision and efficiency, improving the overall system performance.

[0077] The acquisition unit can track the user's eye movements in real time. For example, the acquisition unit can track the user's eye movements in real time using a camera. The acquisition unit can also detect eye movements using an infrared sensor. Furthermore, the acquisition unit can use multiple sensors in combination to acquire the user's eye movements with high accuracy. For example, the acquisition unit can combine a camera and an infrared sensor to track the user's eye movements with high accuracy. This enables more accurate gaze correction by tracking the user's eye movements in real time. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's eye movement data acquired by the camera into a generating AI and have the generating AI perform the processing of tracking in real time.

[0078] The correction unit can correct the eye movements acquired by the acquisition unit to match the camera's gaze. For example, the correction unit uses an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze. The correction unit can also correct the user's eye movements in real time. Furthermore, the correction unit can adjust the timing of the correction to make the user's eye movements appear more natural. For example, the correction unit uses an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze. This allows the user to look the other person in the eye while talking by correcting the eye movements to match the camera's gaze. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input the eye movement data acquired by the acquisition unit into a generating AI and have the generating AI perform the process of correcting it to match the camera's gaze.

[0079] The voice acquisition unit can acquire the user's voice in real time. For example, the voice acquisition unit can acquire the user's voice in real time using a microphone. The voice acquisition unit can also acquire clear audio using noise cancellation technology. Furthermore, the voice acquisition unit can use multiple microphones in combination to acquire the user's voice with high accuracy. For example, the voice acquisition unit can acquire the user's voice with high accuracy by combining a microphone and noise cancellation technology. This enables instant translation and reading aloud by acquiring the user's voice in real time. Some or all of the above processing in the voice acquisition unit may be performed using AI, for example, or without AI. For example, the voice acquisition unit can input the user's voice data acquired by the microphone into a generating AI and have the generating AI execute the process of acquiring the voice in real time.

[0080] The translation unit can translate the audio acquired by the audio acquisition unit and read the translated audio aloud in the user's own voice. For example, the translation unit can use speech recognition technology to convert the user's voice into text and then translate that text. The translation unit can also use speech synthesis technology to read the translated audio aloud in the user's own voice. Furthermore, the translation unit can learn the user's voice and read it aloud in a more natural voice. For example, the translation unit can use speech recognition technology to convert the user's voice into text and then translate that text. By translating the audio and reading it aloud in the user's own voice, language barriers are eliminated, and natural communication is achieved. Some or all of the above processing in the translation unit may be performed using AI, for example, or without AI. For example, the translation unit can input the audio data acquired by the audio acquisition unit into a generating AI and have the generating AI perform the translation and speech synthesis processes.

[0081] The lip correction unit can correct lip movements to match translated speech. For example, the lip correction unit can track the user's lip movements in real time and correct them to match the translated speech. The lip correction unit can also adjust the timing of the corrections to make the user's lip movements appear more natural. Furthermore, the lip correction unit can use a combination of multiple sensors to correct the user's lip movements with high precision. For example, the lip correction unit can combine a camera and an infrared sensor to track the user's lip movements with high precision and correct them to match the translated speech. This eliminates visual inconsistencies by correcting lip movements to match the translated speech. Some or all of the above processing in the lip correction unit may be performed using AI, for example, or without AI. For example, the lip correction unit can input data on the user's lip movements into a generating AI and have the generating AI perform the process of correcting the lip movements in real time.

[0082] The acquisition unit can estimate the user's emotions and adjust the timing of eye movement acquisition based on the estimated emotions. For example, if the user is nervous, the acquisition unit can increase the frequency of eye movement acquisition to collect more detailed data. Conversely, if the user is relaxed, the acquisition unit can decrease the frequency of eye movement acquisition to maintain a natural conversation. Furthermore, if the user is focused, the acquisition unit can prioritize acquiring eye movements for a specific viewpoint. For example, if the user is nervous, the acquisition unit can increase the frequency of eye movement acquisition to collect more detailed data. This allows for more natural communication by adjusting the timing of eye movement acquisition based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input user emotion data into the generative AI and cause the generative AI to perform the process of adjusting the timing of eye movement acquisition.

[0083] The acquisition unit can analyze the user's past eye movement data and select the optimal acquisition method. For example, the acquisition unit can identify points that the user has frequently looked at in the past and prioritize acquiring eye movements to those points. The acquisition unit can also analyze the user's past eye movement data to identify eye movement patterns in specific time periods and adjust the acquisition method. Furthermore, the acquisition unit can predict eye movements under specific circumstances based on the user's past eye movement data and optimize the acquisition method. For example, the acquisition unit can identify points that the user has frequently looked at in the past and prioritize acquiring eye movements to those points. By analyzing past eye movement data, the acquisition unit can select the optimal acquisition method and improve accuracy. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's past eye movement data into a generating AI and have the generating AI perform the process of selecting the optimal acquisition method.

[0084] The acquisition unit can identify the user's current gaze focus when acquiring eye movements and prioritize acquiring gaze towards specific objects. For example, if the user is looking at a specific icon on the screen, the acquisition unit will prioritize acquiring the gaze movement toward that icon. The acquisition unit can also prioritize acquiring the gaze movement toward the face of the person the user is talking to. Furthermore, if the user is reading a document, the acquisition unit can also prioritize acquiring the gaze movement toward that document. For example, if the acquisition unit is looking at a specific icon on the screen, the acquisition unit will prioritize acquiring the gaze movement toward that icon. By identifying the user's gaze focus and prioritizing the acquisition of gaze toward specific objects, more accurate data can be collected. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input the user's gaze focus data into a generating AI and cause the generating AI to execute the process of prioritizing the acquisition of gaze toward specific objects.

[0085] The acquisition unit can estimate the user's emotions and determine the priority of eye movements to acquire based on the estimated user emotions. For example, if the user is tense, the acquisition unit will prioritize acquiring changes in eye movements. It can also prioritize acquiring fixed points of gaze if the user is relaxed. Furthermore, if the user is concentrating, the acquisition unit can prioritize acquiring subtle eye movements. For example, if the acquisition unit is tense, it will prioritize acquiring changes in eye movements. This prioritizes the acquisition of more important data by determining the priority of eye movements based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the acquisition unit may be performed using AI, or not. For example, the acquisition unit can input user emotion data into the generative AI and have the generative AI perform the process of determining the priority of eye movements.

[0086] The acquisition unit can adjust the acquisition accuracy when acquiring eye movements, taking into account changes in the user's ambient light. For example, if the ambient light is bright, the acquisition unit can increase the acquisition accuracy to acquire detailed eye movements. Conversely, if the ambient light is dim, the acquisition unit can also decrease the acquisition accuracy to reduce noise. Furthermore, if the ambient light fluctuates, the acquisition unit can adjust the acquisition accuracy in real time. For example, if the ambient light is bright, the acquisition unit can increase the acquisition accuracy to acquire detailed eye movements. By adjusting the acquisition accuracy to take into account changes in ambient light, more accurate eye movement data can be acquired. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input ambient light data into a generating AI and have the generating AI perform the process of adjusting the acquisition accuracy.

[0087] The acquisition unit can correct the direction of the gaze in conjunction with the user's head movements when acquiring eye movements. For example, the acquisition unit corrects the direction of the gaze in real time when the user moves their head. The acquisition unit can also correct the angle of the gaze to acquire accurate data when the user tilts their head. Furthermore, the acquisition unit can correct the height of the gaze when the user moves their head up and down. For example, the acquisition unit corrects the direction of the gaze in real time when the user moves their head. This allows for the acquisition of more accurate gaze data by correcting the direction of the gaze in conjunction with the user's head movements. Some or all of the above processing in the acquisition unit may be performed using AI, for example, or without AI. For example, the acquisition unit can input data of the user's head movements into a generating AI and have the generating AI perform the process of correcting the direction of the gaze.

[0088] The correction unit can estimate the user's emotions and adjust the method of correcting the camera gaze based on the estimated user emotions. For example, if the user is tense, the correction unit can strengthen the correction of the camera gaze to improve the stability of the gaze. Conversely, if the user is relaxed, the correction unit can also relax the correction of the camera gaze to maintain a natural gaze. Furthermore, if the user is concentrating, the correction unit can prioritize the correction of the camera gaze for a specific viewpoint. For example, if the user is tense, the correction unit can strengthen the correction of the camera gaze to improve the stability of the gaze. This results in a more natural gaze by adjusting the method of correcting the camera gaze based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input user emotion data into the generative AI and cause the generative AI to perform the process of adjusting the method of correcting the camera gaze.

[0089] The correction unit analyzes the user's facial expression during correction and can correct the gaze while maintaining a natural expression. For example, when the user is smiling, the correction unit can maintain the smile while correcting the gaze. The correction unit can also maintain a serious expression when the user is serious while correcting the gaze. Furthermore, the correction unit can maintain a surprised expression when the user is surprised while correcting the gaze. For example, when the user is smiling, the correction unit can maintain the smile while correcting the gaze. In this way, by analyzing the user's facial expression and correcting the gaze while maintaining a natural expression, visual unnaturalness is reduced. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input the user's facial expression data into a generating AI and cause the generating AI to perform the process of correcting the gaze while maintaining a natural expression.

[0090] The correction unit can optimize the timing of correction by taking into account the speed of the user's eye movements. For example, if the user's eye movements are fast, the correction unit can shorten the correction timing to improve real-time performance. Conversely, if the user's eye movements are slow, the correction unit can extend the correction timing to maintain natural movement. Furthermore, if the user's eye movements are irregular, the correction unit can dynamically adjust the correction timing. For example, if the user's eye movements are fast, the correction unit can shorten the correction timing to improve real-time performance. This optimizes the correction timing by taking into account the speed of the user's eye movements, thereby achieving more natural gaze correction. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input user eye movement data into a generating AI and have the generating AI perform the process of optimizing the correction timing.

[0091] The correction unit can estimate the user's emotions and adjust the display method of the corrected gaze based on the estimated user emotions. For example, if the user is nervous, the correction unit provides a simple and highly visible display method. The correction unit can also provide a display method that includes detailed information if the user is relaxed. Furthermore, if the user is in a hurry, the correction unit can provide a concise display method. For example, if the correction unit is nervous, it provides a simple and highly visible display method. This improves visibility by adjusting the display method of the corrected gaze based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input user emotion data into the generative AI and cause the generative AI to perform the process of adjusting the display method of the corrected gaze.

[0092] The correction unit can adjust the correction accuracy during correction, taking into account whether the user is wearing glasses or contact lenses. For example, if the user is wearing glasses, the correction unit can increase the correction accuracy to maintain the accuracy of the gaze. The correction unit can also adjust the correction accuracy to maintain the naturalness of the gaze if the user is wearing contact lenses. Furthermore, if the user is not wearing glasses or contact lenses, the correction unit can optimize the correction accuracy to achieve both accuracy and naturalness of the gaze. For example, if the user is wearing glasses, the correction unit can increase the correction accuracy to maintain the accuracy of the gaze. This maintains the accuracy of the gaze by adjusting the correction accuracy considering whether the user is wearing glasses or contact lenses. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input data on whether the user is wearing glasses or contact lenses into a generating AI and have the generating AI perform the process of adjusting the correction accuracy.

[0093] The correction unit can correct the user's gaze in conjunction with the orientation of their face during correction. For example, when the user moves their face from side to side, the correction unit maintains the orientation of the face while correcting the gaze. The correction unit can also maintain the orientation of the face while correcting the gaze when the user moves their face up and down. Furthermore, the correction unit can maintain the orientation of the face while correcting the gaze when the user moves their face diagonally. For example, when the user moves their face from side to side, the correction unit maintains the orientation of the face while correcting the gaze. This results in a more natural gaze by correcting the gaze in conjunction with the orientation of the user's face. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input data on the orientation of the user's face into a generating AI and have the generating AI execute the process of correcting the gaze.

[0094] The correction unit can correct the user's gaze in conjunction with the orientation of their face during correction. For example, when the user moves their face from side to side, the correction unit maintains the orientation of the face while correcting the gaze. The correction unit can also maintain the orientation of the face while correcting the gaze when the user moves their face up and down. Furthermore, the correction unit can maintain the orientation of the face while correcting the gaze when the user moves their face diagonally. For example, when the user moves their face from side to side, the correction unit maintains the orientation of the face while correcting the gaze. This results in a more natural gaze by correcting the gaze in conjunction with the orientation of the user's face. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can input data on the orientation of the user's face into a generating AI and have the generating AI execute the process of correcting the gaze.

[0095] The voice acquisition unit can estimate the user's emotions and adjust the timing of voice acquisition based on the estimated emotions. For example, if the user is nervous, the voice acquisition unit can increase the frequency of voice acquisition to collect more detailed data. Conversely, if the user is relaxed, the voice acquisition unit can decrease the frequency of voice acquisition to maintain a natural conversation. Furthermore, if the user is focused, the voice acquisition unit can prioritize acquiring voice for specific statements. For example, if the voice acquisition unit is nervous, it can increase the frequency of voice acquisition to collect more detailed data. This allows for a more natural conversation by adjusting the timing of voice acquisition based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the voice acquisition unit may be performed using AI, or not using AI. For example, the voice acquisition unit can input user emotion data into the generative AI and have the generative AI perform the process of adjusting the timing of voice acquisition.

[0096] The voice acquisition unit can analyze the user's past voice data and select the optimal acquisition method. For example, the voice acquisition unit can identify phrases that the user has frequently uttered in the past and prioritize acquiring audio for those phrases. The voice acquisition unit can also analyze the user's past voice data to identify patterns in speech during specific time periods and adjust the acquisition method accordingly. Furthermore, the voice acquisition unit can predict speech under specific circumstances based on the user's past voice data and optimize the acquisition method. For example, the voice acquisition unit can identify phrases that the user has frequently uttered in the past and prioritize acquiring audio for those phrases. By analyzing past voice data, the optimal acquisition method is selected, improving accuracy. Some or all of the above processing in the voice acquisition unit may be performed using AI, for example, or without AI. For example, the voice acquisition unit can input the user's past voice data into a generating AI and have the generating AI perform the process of selecting the optimal acquisition method.

[0097] The voice acquisition unit can acquire clear audio by filtering out ambient noise around the user during voice acquisition. For example, if the user is in a noisy environment, the voice acquisition unit can acquire clear audio using noise cancellation technology. Furthermore, if the user is in a quiet environment, the voice acquisition unit can acquire natural audio with minimal noise filtering. Additionally, if the user is moving, the voice acquisition unit can acquire clear audio by filtering out ambient noise in real time. For example, if the user is in a noisy environment, the voice acquisition unit can acquire clear audio using noise cancellation technology. This filters out ambient noise, resulting in clear audio. Some or all of the above processing in the voice acquisition unit may be performed using AI, or without AI. For example, the voice acquisition unit can input ambient noise data into a generating AI and have the generating AI perform the noise filtering process.

[0098] The voice acquisition unit can estimate the user's emotions and determine the priority of the audio to acquire based on the estimated user emotions. For example, if the user is nervous, the voice acquisition unit will prioritize acquiring important statements. It can also prioritize acquiring natural conversational flows if the user is relaxed. Furthermore, if the user is focused, the voice acquisition unit can prioritize acquiring statements on specific topics. For example, if the voice acquisition unit is nervous, it will prioritize acquiring important statements. This prioritizes the acquisition of important statements by determining the audio priority based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the voice acquisition unit may be performed using AI, or not. For example, the voice acquisition unit can input user emotion data into the generative AI and have the generative AI perform the process of determining the audio priority.

[0099] The audio acquisition unit can adjust the acquisition accuracy when acquiring audio, taking into account the user's microphone position. For example, if the user places the microphone close by, the audio acquisition unit can increase the acquisition accuracy to acquire detailed audio. Conversely, if the user places the microphone far away, the audio acquisition unit can also adjust the acquisition accuracy to acquire clear audio. Furthermore, if the user moves the microphone, the audio acquisition unit can adjust the acquisition accuracy in real time. For example, if the user places the microphone close by, the audio acquisition unit can increase the acquisition accuracy to acquire detailed audio. This ensures clear audio is acquired by adjusting the acquisition accuracy considering the microphone position. Some or all of the above processing in the audio acquisition unit may be performed using AI, for example, or without AI. For example, the audio acquisition unit can input microphone position data into a generating AI and have the generating AI perform the process of adjusting the acquisition accuracy.

[0100] The voice acquisition unit can optimize its voice acquisition method according to the user's speaking speed. For example, if the user speaks quickly, the voice acquisition unit adjusts the acquisition method to improve real-time performance. The voice acquisition unit can also adjust the acquisition method to acquire more detailed audio if the user speaks slowly. Furthermore, if the user speaks at an irregular speed, the voice acquisition unit can dynamically adjust the acquisition method to acquire the optimal audio. For example, if the user speaks quickly, the voice acquisition unit adjusts the acquisition method to improve real-time performance. This optimizes the acquisition method according to the user's speaking speed, resulting in the acquisition of more accurate audio data. Some or all of the above-described processes in the voice acquisition unit may be performed using AI, for example, or without AI. For example, the voice acquisition unit can input user speaking speed data into a generating AI and have the generating AI perform the process of optimizing the acquisition method.

[0101] The translation unit can estimate the user's emotions and adjust the translation's expression based on the estimated emotions. For example, if the user is nervous, the translation unit will provide a simple, literal translation. If the user is relaxed, the translation unit can also provide a translation with more natural expressions. Furthermore, if the user is excited, the translation unit can provide an emotionally emphasized translation. For example, if the user is nervous, the translation unit will provide a simple, literal translation. By adjusting the translation's expression based on the user's emotions, a more natural translation can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the translation unit may be performed using AI, for example, or not using AI. For example, the translation unit can input user emotion data into the generative AI and have the generative AI perform the process of adjusting the translation's expression.

[0102] The translation unit can improve translation accuracy by taking into account the user's specialized terminology and slang during translation. For example, if the user frequently uses specialized terminology, the translation unit will accurately translate that terminology. The translation unit can also appropriately translate slang if the user uses it. Furthermore, if the user uses specific industry-specific terminology, the translation unit can accurately translate that industry-specific terminology. For example, if the user frequently uses specialized terminology, the translation unit will accurately translate that terminology. This improves translation accuracy by considering specialized terminology and slang. Some or all of the above processing in the translation unit may be performed using AI, for example, or without AI. For example, the translation unit can input data on the user's specialized terminology and slang into a generating AI and have the generating AI perform processing to improve translation accuracy.

[0103] The translation unit can optimize the timing of translation according to the user's speaking speed. For example, if the user speaks quickly, the translation unit can shorten the translation timing to improve real-time performance. Conversely, if the user speaks slowly, the translation unit can extend the translation timing to provide a more detailed translation. Furthermore, the translation unit can dynamically adjust the translation timing if the user speaks at an irregular pace. For example, if the user speaks quickly, the translation unit can shorten the translation timing to improve real-time performance. This improves real-time performance by optimizing the translation timing according to the user's speaking speed. Some or all of the above processing in the translation unit may be performed using AI, for example, or without AI. For example, the translation unit can input user speaking speed data into a generating AI and have the generating AI perform the process of optimizing the translation timing.

[0104] The translation unit can estimate the user's emotions and adjust the length of the translation based on the estimated emotions. For example, if the user is nervous, the translation unit will provide a short, to-the-point translation. If the user is relaxed, the translation unit may also provide a longer translation that includes more detailed explanations. Furthermore, if the user is excited, the translation unit may also provide a longer translation that emphasizes the emotion. For example, if the user is nervous, the translation unit will provide a short, to-the-point translation. By adjusting the length of the translation based on the user's emotions, a more appropriate translation can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the translation unit may be performed using AI, for example, or not using AI. For example, the translation unit can input user emotion data into the generative AI and have the generative AI perform the process of adjusting the length of the translation.

[0105] The translation unit can adjust translation accuracy by taking into account regional differences in the user's spoken language during translation. For example, if the user speaks a dialect from a specific region, the translation unit will translate that dialect appropriately. The translation unit can also take into account regional differences if the user speaks languages ​​from different regions. Furthermore, if the user speaks a mix of languages ​​from multiple regions, the translation unit can also take into account regional differences. For example, if the translation unit speaks a dialect from a specific region, the translation unit will translate that dialect appropriately. This provides a more accurate translation by taking regional differences into account. Some or all of the above processing in the translation unit may be performed using AI, for example, or not using AI. For example, the translation unit can input regional difference data of the user's spoken language into a generating AI and have the generating AI perform the process of adjusting translation accuracy.

[0106] The translation unit can apply different translation algorithms depending on the category of what the user is saying during translation. For example, if the user uses business terminology, the translation unit will apply a business translation algorithm. It can also apply a conversational translation algorithm if the user is engaging in everyday conversation. Furthermore, if the user uses technical terminology, the translation unit can apply a technical translation algorithm. This improves translation accuracy by applying translation algorithms appropriate to the content category. Some or all of the above processing in the translation unit may be performed using AI, for example, or without AI. For example, the translation unit can input category data of what the user is saying into a generating AI and have the generating AI perform the process of applying different translation algorithms.

[0107] The lip correction unit can estimate the user's emotions and adjust the method of correcting lip movements based on the estimated emotions. For example, if the user is nervous, the lip correction unit can emphasize lip movements to reduce visual discomfort. It can also correct lip movements to keep them natural if the user is relaxed. Furthermore, if the user is excited, the lip correction unit can dynamically correct lip movements. For example, if the user is nervous, the lip correction unit can emphasize lip movements to reduce visual discomfort. This reduces visual discomfort by adjusting the method of correcting lip movements based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the lip correction unit may be performed using AI, or not. For example, the lip correction unit can input user emotion data into the generative AI and have the generative AI perform the process of adjusting the method of correcting lip movements.

[0108] The lip correction unit analyzes the user's facial expression during correction and can correct lip movements while maintaining a natural expression. For example, when the user is smiling, the lip correction unit can maintain the smile by correcting lip movements. It can also maintain a serious expression when the user is serious by correcting lip movements. Furthermore, it can maintain a surprised expression when the user is surprised by correcting lip movements. For example, when the user is smiling, the lip correction unit can maintain the smile by correcting lip movements. This reduces visual unnaturalness by analyzing the user's facial expression and correcting lip movements while maintaining a natural expression. Some or all of the above processing in the lip correction unit may be performed using AI, for example, or without AI. For example, the lip correction unit can input the user's facial expression data into a generating AI and cause the generating AI to perform the process of correcting lip movements while maintaining a natural expression.

[0109] The lip correction unit can optimize the timing of lip movement correction according to the user's speaking speed during correction. For example, if the user speaks quickly, the lip correction unit can shorten the timing of lip movement correction to improve real-time performance. Conversely, if the user speaks slowly, the lip correction unit can extend the timing of lip movement correction to maintain natural movement. Furthermore, if the user speaks at an irregular speed, the lip correction unit can dynamically adjust the timing of lip movement correction. For example, if the user speaks quickly, the lip correction unit can shorten the timing of lip movement correction to improve real-time performance. This optimizes the timing of lip movement correction according to the user's speaking speed, resulting in more natural movement. Some or all of the above processing in the lip correction unit may be performed using AI, for example, or without AI. For example, the lip correction unit can input user speaking speed data into a generating AI and have the generating AI perform the process of optimizing the timing of lip movement correction.

[0110] The lip correction unit can estimate the user's emotions and adjust the display method of the corrected lip movements based on the estimated emotions. For example, if the user is nervous, the lip correction unit can provide a simple and highly visible display method. It can also provide a display method that includes detailed information if the user is relaxed. Furthermore, if the user is in a hurry, it can provide a concise display method. For example, if the lip correction unit is nervous, it provides a simple and highly visible display method. This improves visibility by adjusting the display method of the corrected lip movements based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the lip correction unit may be performed using AI, for example, or without AI. For example, the lip correction unit can input user emotion data into the generative AI and have the generative AI perform the process of adjusting the display method of the corrected lip movements.

[0111] The lip correction unit can correct lip movement in conjunction with the user's facial orientation during correction. For example, when the user moves their face from side to side, the lip correction unit can correct lip movement while maintaining the facial orientation. It can also correct lip movement while maintaining the facial orientation when the user moves their face up and down. Furthermore, it can correct lip movement while maintaining the facial orientation when the user moves their face diagonally. For example, when the user moves their face from side to side, the lip correction unit corrects lip movement while maintaining the facial orientation. This allows for more natural movement by correcting lip movement in conjunction with the user's facial orientation. Some or all of the above-described processes in the lip correction unit may be performed using AI, or without AI. For example, the lip correction unit can input user facial orientation data into a generating AI and have the generating AI perform the process of correcting lip movement.

[0112] The lip correction unit can apply different lip movement correction algorithms depending on the category of what the user is saying during correction. For example, if the user is using business terminology, the lip correction unit can apply a business-oriented lip movement correction algorithm. It can also apply a lip movement correction algorithm for everyday conversation if the user is engaging in everyday conversation. Furthermore, if the user is using technical terminology, the lip correction unit can apply a technical-oriented lip movement correction algorithm. For example, if the user is using business terminology, the lip correction unit can apply a business-oriented lip movement correction algorithm. This allows for more natural movement by applying a lip movement correction algorithm appropriate to the content category. Some or all of the above processing in the lip correction unit may be performed using AI, or without AI. For example, the lip correction unit can input category data of what the user is saying into a generating AI and have the generating AI perform the process of applying different lip movement correction algorithms.

[0113] The lip correction unit can optimize the method of correcting lip movements according to the user's speaking speed during correction. For example, if the user speaks quickly, the lip correction unit adjusts the method of correcting lip movements to improve real-time performance. The lip correction unit can also adjust the method of correcting lip movements to maintain natural movement if the user speaks slowly. Furthermore, the lip correction unit can dynamically adjust the method of correcting lip movements if the user speaks at an irregular speed. For example, if the user speaks quickly, the lip correction unit adjusts the method of correcting lip movements to improve real-time performance. This optimizes the method of correcting lip movements according to the user's speaking speed, resulting in more natural movement. Some or all of the above processing in the lip correction unit may be performed using AI, for example, or without AI. For example, the lip correction unit can input user speaking speed data into a generating AI and have the generating AI perform the process of optimizing the method of correcting lip movements.

[0114] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0115] Next-generation video conferencing systems can also include a background recognition unit that automatically recognizes the user's background and switches to an appropriate virtual background. For example, if the user is at home, the background recognition unit can switch to a virtual office background. It can also switch to a virtual cafe background if the user is out and about. Furthermore, if the user is in a meeting, the background recognition unit can switch to a virtual conference room background. This allows for a more professional impression by automatically recognizing the user's background and switching to an appropriate virtual background.

[0116] Next-generation video call systems can also be equipped with an emotion feedback unit that analyzes the user's facial expressions in real time and provides emotion-based feedback. For example, the emotion feedback unit can display positive feedback when the user smiles. It can also display feedback to encourage concentration when the user has a serious expression. Furthermore, it can display feedback to share the surprise when the user has a surprised expression. This improves the quality of communication by analyzing the user's facial expressions in real time and providing emotion-based feedback that matches their feelings.

[0117] Next-generation video call systems can also include a voice feedback unit that analyzes the user's voice tone and provides emotionally responsive voice feedback. For example, if the user is tense, the voice feedback unit can provide relaxing voice feedback. It can also provide positive voice feedback if the user is relaxed. Furthermore, if the user is excited, it can provide calming voice feedback. This allows for more natural communication by analyzing the user's voice tone and providing emotionally responsive voice feedback.

[0118] Next-generation video conferencing systems can also include a gesture recognition unit that recognizes user gestures and performs actions corresponding to those gestures. For example, if a user raises their hand, the gesture recognition unit can notify the system of a request to speak. It can also highlight a specific object if a user points. Furthermore, if a user claps, the gesture recognition unit can display positive feedback. This allows for more interactive communication by recognizing user gestures and performing actions accordingly.

[0119] Next-generation video conferencing systems can also include an information provision unit that analyzes user speech in real time and provides relevant information. For example, if a user is discussing a specific topic, the information provision unit can display news articles related to that topic. It can also provide answers to questions asked by the user. Furthermore, if a user is referring to specific data, the information provision unit can display details of that data. This allows for improved conversation quality by analyzing user speech in real time and providing relevant information.

[0120] Next-generation video call systems can also include a conversation tone adjustment unit that estimates the user's emotions and adjusts the tone of the conversation based on those emotions. For example, if the user is nervous, the conversation tone adjustment unit will conduct the conversation in a relaxed tone. It can also conduct the conversation in a positive tone if the user is relaxed. Furthermore, if the user is excited, the conversation tone adjustment unit can conduct the conversation in a calm tone. This allows for more natural communication by adjusting the tone of the conversation based on the user's emotions.

[0121] Next-generation video conferencing systems can also include an intent estimation unit that analyzes user speech and estimates the intent behind their statements. For example, if a user asks a question, the intent estimation unit can estimate the intent behind that question and provide an appropriate answer. Furthermore, if a user makes a suggestion, the intent estimation unit can estimate the intent behind that suggestion and provide relevant information. Additionally, if a user expresses an opinion, the intent estimation unit can estimate the intent behind that opinion and provide appropriate feedback. This allows for more effective communication by analyzing user speech and estimating their intent.

[0122] Next-generation video call systems can also include a conversation support unit that estimates the user's emotions and supports the flow of the conversation based on those emotions. For example, if the user is nervous, the conversation support unit can offer topics that promote relaxation. It can also offer positive topics if the user is relaxed. Furthermore, if the user is excited, it can offer topics that promote calmness. This allows for more natural communication by supporting the flow of the conversation based on the user's emotions.

[0123] Next-generation video conferencing systems can also include a text display unit that translates user speech in real time and displays the translation as text. For example, if a user is speaking in English, the text display unit will translate their speech into Japanese and display it as text. It can also translate a user's speech into English if they are speaking in Japanese and display it as text. Furthermore, if a user is using multiple languages, the text display unit can translate their speech into each language and display it as text. This allows for real-time translation of user speech and display of the translation as text, eliminating language barriers and facilitating smoother communication.

[0124] Next-generation video call systems can also include a conversation content adjustment unit that estimates the user's emotions and adjusts the conversation content based on those emotions. For example, if the user is nervous, the conversation content adjustment unit can adjust the content to promote relaxation. It can also adjust the content to be positive if the user is relaxed. Furthermore, if the user is excited, the conversation content adjustment unit can adjust the content to promote calmness. By adjusting the conversation content based on the user's emotions, more natural communication can be achieved.

[0125] The following briefly describes the processing flow for example form 2.

[0126] Step 1: The acquisition unit acquires the user's eye movements. The acquisition unit can, for example, track the user's eye movements in real time using a camera. It can also detect eye movements using an infrared sensor. Furthermore, multiple sensors can be used in combination. For example, a camera and an infrared sensor can be combined to track the user's eye movements with high precision. Step 2: The correction unit corrects the eye movements acquired by the acquisition unit to match the camera's gaze. The correction unit uses an algorithm that calculates the angle of gaze and corrects it to match the camera's gaze, for example. It can also correct the user's eye movements in real time. Furthermore, the timing of the correction can be adjusted. Step 3: The voice acquisition unit acquires the user's voice. The voice acquisition unit can acquire the user's voice in real time, for example, using a microphone. It can also acquire clear audio using noise cancellation technology. Furthermore, multiple microphones can be used in combination. For example, a microphone and noise cancellation technology can be combined to acquire the user's voice with high accuracy. Step 4: The translation unit translates the audio acquired by the audio acquisition unit and reads the translated audio aloud in the user's voice. The translation unit can, for example, use speech recognition technology to convert the user's voice into text and then translate that text. It can also use speech synthesis technology to read the translated audio aloud in the user's voice. Furthermore, it can learn the user's voice and read aloud in a more natural voice. Step 5: The lip correction unit corrects lip movements to match the translated audio. For example, the lip correction unit tracks the user's lip movements in real time and corrects them to match the translated audio. The timing of the correction can also be adjusted. Furthermore, multiple sensors can be used in combination. For example, a camera and an infrared sensor can be combined to track the user's lip movements with high precision and correct them to match the translated audio.

[0127] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0128] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0129] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0130] Each of the multiple elements described above, including the acquisition unit, correction unit, voice acquisition unit, translation unit, and lip correction unit, is implemented, for example, in at least one of the smart device 14 and the data processing unit 12. For example, the acquisition unit tracks the user's eye movements in real time using the camera 42 and infrared sensor of the smart device 14. The correction unit uses an algorithm that calculates the angle of gaze using the specific processing unit 290 of the data processing unit 12 and corrects it to match the camera's gaze. The voice acquisition unit acquires the user's voice in real time using the microphone 38B of the smart device 14. The translation unit converts the user's voice into text using speech recognition technology using the specific processing unit 290 of the data processing unit 12 and translates that text. The lip correction unit tracks the user's lip movements in real time using the camera 42 and infrared sensor of the smart device 14 and corrects them to match the translated voice. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0131] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0132] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0133] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0134] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0135] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0136] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0137] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0138] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0139] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0140] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0141] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0142] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0143] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0144] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0145] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0146] Each of the multiple elements described above, including the acquisition unit, correction unit, voice acquisition unit, translation unit, and lip correction unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the acquisition unit tracks the user's eye movements in real time using the camera 42 and infrared sensor of the smart glasses 214. The correction unit calculates the angle of gaze using the identification processing unit 290 of the data processing unit 12 and uses an algorithm to correct it to the camera's line of sight. The voice acquisition unit acquires the user's voice in real time using the microphone 238 of the smart glasses 214. The translation unit converts the user's voice into text using speech recognition technology with the identification processing unit 290 of the data processing unit 12 and translates that text. The lip correction unit tracks the user's lip movements in real time using the camera 42 and infrared sensor of the smart glasses 214 and corrects them to match the translated voice. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0147] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0148] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0149] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0150] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0151] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0152] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0153] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0154] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0155] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0156] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0157] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0158] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0159] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0160] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0161] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0162] Each of the multiple elements described above, including the acquisition unit, correction unit, voice acquisition unit, translation unit, and lip correction unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the acquisition unit tracks the user's eye movements in real time using the camera 42 and infrared sensor of the headset terminal 314. The correction unit uses an algorithm that calculates the angle of gaze using the specific processing unit 290 of the data processing unit 12 and corrects it to match the camera's gaze. The voice acquisition unit acquires the user's voice in real time using the microphone 238 of the headset terminal 314. The translation unit converts the user's voice into text using speech recognition technology using the specific processing unit 290 of the data processing unit 12 and translates that text. The lip correction unit tracks the user's lip movements in real time using the camera 42 and infrared sensor of the headset terminal 314 and corrects them to match the translated voice. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0163] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0164] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0165] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0166] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0167] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0168] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0169] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0170] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0171] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0172] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0173] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0174] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0175] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0176] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0177] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0178] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0179] Each of the multiple elements described above, including the acquisition unit, correction unit, voice acquisition unit, translation unit, and lip correction unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the acquisition unit tracks the user's eye movements in real time using the camera 42 and infrared sensor of the robot 414. The correction unit uses an algorithm that calculates the angle of gaze using the specific processing unit 290 of the data processing unit 12 and corrects it to the camera's line of sight. The voice acquisition unit acquires the user's voice in real time using the microphone 238 of the robot 414. The translation unit converts the user's voice into text using speech recognition technology using the specific processing unit 290 of the data processing unit 12 and translates that text. The lip correction unit tracks the user's lip movements in real time using the camera 42 and infrared sensor of the robot 414 and corrects them to match the translated voice. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0180] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0181] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0182] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0183] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0184] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0185] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0186] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0187] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0188] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0189] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0190] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0191] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0192] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0193] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0194] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0195] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0196] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0197] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0198] (Note 1) An acquisition unit that acquires the user's eye movements, A correction unit corrects the eye movements acquired by the acquisition unit to match the camera's gaze, A sound acquisition unit that acquires sound, A translation unit that translates the audio acquired by the aforementioned audio acquisition unit and reads the translated audio aloud in the person's voice, It includes a lip correction unit that corrects lip movements to match the translated audio. A system characterized by the following features. (Note 2) The acquisition unit is, Track the user's eye movements in real time. The system described in Appendix 1, characterized by the features described herein. (Note 3) The correction unit, Based on the eye movements acquired by the aforementioned acquisition unit, the camera's gaze is corrected. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned audio acquisition unit, Capture user voice in real time. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned translation department, The voice acquisition unit acquires the voice and translates it, and the translated voice is read aloud in the person's own voice. The system described in Appendix 1, characterized by the features described herein. (Note 6) The lip correction portion is, Adjust lip movements to match the translated audio. The system described in Appendix 1, characterized by the features described herein. (Note 7) The acquisition unit is, The system estimates the user's emotions and adjusts the timing of eye movement acquisition based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The acquisition unit is, Analyze the user's past eye movement data and select the optimal acquisition method. The system described in Appendix 1, characterized by the features described herein. (Note 9) The acquisition unit is, When capturing eye movements, the system identifies the user's current gaze focus and prioritizes capturing gazes directed at specific objects. The system described in Appendix 1, characterized by the features described herein. (Note 10) The acquisition unit is, It estimates the user's emotions and determines the priority of eye movements to capture based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The acquisition unit is, When acquiring eye movements, the acquisition accuracy is adjusted to take into account changes in the user's ambient light. The system described in Appendix 1, characterized by the features described herein. (Note 12) The acquisition unit is, When acquiring eye movements, the direction of the gaze is corrected in conjunction with the user's head movements. The system described in Appendix 1, characterized by the features described herein. (Note 13) The correction unit, The system estimates the user's emotions and adjusts the camera's gaze correction method based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The correction unit, During correction, the system analyzes the user's facial expression and corrects their gaze while maintaining a natural expression. The system described in Appendix 1, characterized by the features described herein. (Note 15) The correction unit, During calibration, the timing of the calibration is optimized by taking into account the speed of the user's eye movements. The system described in Appendix 1, characterized by the features described herein. (Note 16) The correction unit, The system estimates the user's emotions and adjusts the display method of the corrected gaze based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The correction unit, During calibration, the calibration accuracy is adjusted taking into account whether the user is wearing glasses or contact lenses. The system described in Appendix 1, characterized by the features described herein. (Note 18) The correction unit, During correction, the gaze direction is corrected in conjunction with the user's facial orientation. The system described in Appendix 1, characterized by the features described herein. (Note 19) The correction unit, During correction, the gaze direction is corrected in conjunction with the user's facial orientation. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned audio acquisition unit, The system estimates the user's emotions and adjusts the timing of voice acquisition based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned audio acquisition unit, Analyze the user's past voice data and select the optimal acquisition method. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned audio acquisition unit, When acquiring audio, the system filters out ambient noise around the user to obtain clear audio. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned audio acquisition unit, It estimates the user's emotions and determines the priority of audio to acquire based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned audio acquisition unit, When acquiring audio, the system adjusts the acquisition accuracy by taking into account the user's microphone position. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned audio acquisition unit, When acquiring audio, the acquisition method is optimized according to the user's speaking speed. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned translation department, It estimates the user's emotions and adjusts the translation's expression based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned translation department, During translation, we improve translation accuracy by taking into account the user's specialized terminology and slang. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned translation department, During translation, the timing of the translation is optimized according to the user's speaking speed. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned translation department, It estimates the user's sentiment and adjusts the translation length based on the estimated sentiment. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned translation department, During translation, the system adjusts translation accuracy by taking into account regional differences in the user's spoken language. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned translation department, During translation, different translation algorithms are applied depending on the category of what the user is saying. The system described in Appendix 1, characterized by the features described herein. (Note 32) The lip correction portion is, It estimates the user's emotions and adjusts how lip movements are modified based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 33) The lip correction portion is, During the editing process, the system analyzes the user's facial expressions and corrects lip movements while maintaining a natural expression. The system described in Appendix 1, characterized by the features described herein. (Note 34) The lip correction portion is, During correction, the timing of lip movement corrections is optimized according to the user's speaking speed. The system described in Appendix 1, characterized by the features described herein. (Note 35) The lip correction portion is, It estimates the user's emotions and adjusts how the corrected lip movements are displayed based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 36) The lip correction portion is, During editing, the lip movements are corrected in conjunction with the user's facial orientation. The system described in Appendix 1, characterized by the features described herein. (Note 37) The lip correction portion is, During correction, different lip movement correction algorithms are applied depending on the category of what the user is saying. The system described in Appendix 1, characterized by the features described herein. (Note 38) The lip correction portion is, During editing, the method of correcting lip movements is optimized according to the user's speaking speed. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]

[0199] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. An acquisition unit that acquires the user's eye movements, A correction unit corrects the eye movements acquired by the acquisition unit to match the camera's gaze, A sound acquisition unit that acquires sound, A translation unit that translates the audio acquired by the aforementioned audio acquisition unit and reads the translated audio aloud in the person's voice, It includes a lip correction unit that corrects lip movements to match the translated audio. A system characterized by the following features.

2. The acquisition unit is, Track the user's eye movements in real time. The system according to feature 1.

3. The correction unit, Based on the eye movements acquired by the aforementioned acquisition unit, the camera's gaze is corrected. The system according to feature 1.

4. The aforementioned audio acquisition unit, Capture user voice in real time. The system according to feature 1.

5. The aforementioned translation department, The voice acquisition unit acquires the voice and translates it, and the translated voice is read aloud in the person's own voice. The system according to feature 1.

6. The lip correction portion is, Adjust lip movements to match the translated audio. The system according to feature 1.

7. The acquisition unit is, The system estimates the user's emotions and adjusts the timing of eye movement acquisition based on the estimated emotions. The system according to feature 1.

8. The acquisition unit is, Analyze the user's past eye movement data and select the optimal acquisition method. The system according to feature 1.

9. The acquisition unit is, When capturing eye movements, the system identifies the user's current gaze focus and prioritizes capturing gazes directed at specific objects. The system according to feature 1.

10. The acquisition unit is, It estimates the user's emotions and determines the priority of eye movements to capture based on the estimated user emotions. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A