Control method for image shake correction in a system that enables activities using a doppelganger.
The system addresses image shake issues in video streaming by performing shake correction on the user's terminal device, ensuring stable viewing, battery conservation, and optimized processing, thus improving user experience and efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TORARU CO LTD
- Filing Date
- 2022-03-31
- Publication Date
- 2026-04-21
AI Technical Summary
Existing video streaming technologies fail to adequately correct image shake, leading to user discomfort and motion sickness when viewing videos streamed from a moving substitute device, and may increase battery consumption and processing load.
A system comprising a first terminal device for the user and a second terminal device for the substitute, where image shake correction processing is performed on the user's device to reduce shake, optionally with control and status sharing, and battery management features.
Ensures stable video viewing with reduced shake, conserves battery life, and optimizes processing load by performing image shake correction on the user's device, enhancing user convenience and real-time performance.
Smart Images

Figure 0007849003000001 
Figure 0007849003000002 
Figure 0007849003000003
Abstract
Description
Technical Field
[0001] The present invention relates to a technology for assisting communication using a terminal device.
Background Art
[0002] When one has to go to a certain place (hereinafter referred to as "the site"), it may be helpful to have a person or a device such as a robot (hereinafter referred to as "a substitute") that can go to the site on one's behalf.
[0003] As a patent document that discloses a technology for satisfying the above needs, for example, there is Patent Document 1. In Patent Document 1, a mechanism is proposed that enables a person at a location other than the site to experience as if being at the site by streaming video from a terminal device used by the substitute at the site to the terminal device used by the person via a network such as the Internet.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] When experiencing as if being at the site using the mechanism as described in Patent Document 1, the video distributed while the substitute is moving may shake along with the shaking of the camera, and the person viewing the video may not be able to easily grasp the objects shown in the video or may develop motion sickness.
[0006] One technology that mitigates the above problems is image stabilization, also known as "camera shake correction." Currently, there are two main types of image stabilization methods: optical, which physically adjusts the optical axis, and electronic, which corrects the image data received from the light-receiving element.
[0007] The terminal device used by the avatar may not have a video shake correction function, or even if it does, the function may not be sufficient to reduce video shake. Furthermore, even if the terminal device used by the avatar has a video shake correction function, the avatar may forget to use it or make a mistake in the settings for using the video shake correction function, resulting in the video shake not being reduced.
[0008] In view of the above circumstances, the present invention provides a means to ensure that when a person viewing video streamed from a location is located outside of that location, they can reliably view video with reduced image shake. [Means for solving the problem]
[0009] The present invention relates to a system comprising a first terminal device used by the person concerned, and a second terminal device used by a doppelganger, which is a person or device present at the location acting as a substitute for the person concerned. ,before The second terminal device streams video data representing the video generated by the camera carried by the avatar to the first terminal device via a communication network. The aforementioned system By the first terminal device described above That The present invention provides a method comprising the steps of: applying image shake correction processing, which is data processing that reduces image shake contained in the image represented by the image shake data, to image data streamed from the first terminal device via the communication network, and displaying the image represented by the image shake correction processing on a display. [Effects of the Invention]
[0010] According to the present invention, even if the video streamed from the location contains video shake, the user can reliably view video with reduced shake because the user's terminal device performs video shake correction. [Brief explanation of the drawing]
[0011] [Figure 1] A diagram showing the configuration of a system according to the first embodiment of the present invention. [Figure 2] A diagram showing the configuration of a system according to one modified example of the present invention. [Figure 3] A diagram showing the configuration of a system according to one modified example of the present invention. [Figure 4] A diagram showing the configuration of a system according to one modified example of the present invention. [Modes for carrying out the invention]
[0012] [1. First Embodiment] A first embodiment of the present invention will be described below. Figure 1 is a diagram showing the configuration of System 1 according to the first embodiment of the present invention. System 1 comprises a terminal device 11 used by the person X and a terminal device 12 used by a doppelganger Y located at location S as a substitute for the person X. The doppelganger Y may be a person or a device such as a robot. Terminal devices 11 and 12 communicate data with each other via a communication network 9.
[0013] Terminal devices 11 and 12 are data processing devices equipped with communication functions, such as computers. The type of computer used as terminal device 11 may be any of the following: a desktop PC (Personal Computer), a notebook (or laptop) PC, or a tablet PC (including a smartphone with the function to make calls using a telephone number via a mobile communication network). Terminal device 11 may or may not be portable by user X.
[0014] The type of computer used as the terminal device 12 may be any type such as a notebook (or laptop) PC, a tablet PC (including a smartphone having a function of making a call using a telephone number via a mobile communication network), etc. The terminal device 12 needs to be portable by the avatar Y. Therefore, the terminal device 12 has a built-in battery and operates on the power supplied from the battery.
[0015] For example, the following devices are connected to the terminal device 11. Display An operation device (such as a keyboard, a mouse, a touch panel, etc.) that receives the operation of the person X Speaker Microphone
[0016] In addition, each of the devices connected to the terminal device 11 may be connected to the terminal device 11 as an external device of the terminal device 11, or may be arranged (built-in) inside the housing of the terminal device 11. Hereinafter, a device connected to the terminal device 11 (for example, a display) will be described as a device of the terminal device 11 (for example, the display of the terminal device 11).
[0017] For example, the following devices are connected to the terminal device 12. Display An operation device (such as a keyboard, a mouse, a touch panel, etc.) that receives the operation of the avatar Y Speaker Microphone Camera
[0018] In addition, each of the devices connected to the terminal device 12 may be connected to the terminal device 12 as an external device of the terminal device 12, or may be arranged (built-in) inside the housing of the terminal device 12. Hereinafter, a device connected to the terminal device 12 (for example, a display) will be described as a device of the terminal device 12 (for example, the display of the terminal device 12).
[0019] The avatar Y moves around, for example, at location S, while operating the camera of terminal device 12. During this time, terminal device 12 streams video data representing the image generated by the camera to terminal device 11.
[0020] Terminal device 11 applies image shake correction processing to the video data streamed from terminal device 12. Here, the image shake correction processing is a known electronic image shake correction processing (also called camera shake correction processing). Specifically, electronic image shake correction processing is a process that moves the position of each still image, which is arranged in chronological order and constitutes the video represented by the video data, within the display screen so that high-frequency components of changes in the position of objects depicted in those still images are cut off.
[0021] The terminal device 11 outputs video data that has undergone image shake correction processing to its display, and displays the image represented by the video data on the display.
[0022] As a result, person X can view the video with reduced image shake.
[0023] In System 1, image shake correction processing is performed in terminal device 11. Therefore, regardless of whether or not image shake correction processing is performed in terminal device 12, user X can reliably view images with reduced image shake.
[0024] By the way, image stabilization processing requires power. Therefore, if image stabilization processing is performed in terminal device 12, the battery consumption of terminal device 12 will increase accordingly. In system 1, terminal device 12 does not necessarily have to perform image stabilization processing. Therefore, in system 1, if terminal device 12 has an image stabilization function, the battery consumption of terminal device 12 can be slowed down by not using that image stabilization function.
[0025] Furthermore, image shake correction processing increases the processing load on the processor that performs data processing. Therefore, if image shake correction processing is performed in terminal device 12, the processing load on the processor of terminal device 12 will increase accordingly. As previously mentioned, in system 1, terminal device 12 does not necessarily have to perform image shake correction processing. Therefore, in system 1, if terminal device 12 has an image shake correction function, the processing load on the processor of terminal device 12 can be reduced by not using that image shake correction function.
[0026] Furthermore, there are cases where the processing power of the processor in terminal device 11 is higher than that of terminal device 12. In that case, with system 1, where video shake correction processing is performed in terminal device 11, the timing at which the video is displayed to user X is faster, and real-time performance is improved compared to a normal system where video shake correction processing is performed in terminal device 12.
[0027] The video shake correction processing performed by the terminal device 11 may be carried out solely by the processor of the terminal device 11, or it may be carried out in cooperation with the processor of the terminal device 11 and a GPU (Graphics Processing Unit) connected to the terminal device 11 as an external device.
[0028] [2. Second Embodiment] A second embodiment of the present invention will be described below. The system configuration according to the second embodiment is the same as that of system 1 according to the first embodiment shown in Figure 1. The differences between the second embodiment and the first embodiment will be described below, while the common points will not be explained.
[0029] In System 1 according to the second embodiment (hereinafter simply referred to as "System 1"), the terminal device 12 is equipped with a video shake correction function. The video shake correction function provided by the terminal device 12 may be either optical or electronic.
[0030] The terminal device 12 then performs control related to the video shake correction process in accordance with instructions transmitted from the terminal device 11. Control related to the video shake correction process includes, for example, starting or ending the video shake correction process, and changing settings related to the video shake correction process (for example, changing the correction intensity).
[0031] Terminal device 12 continuously transmits status data to terminal device 11 indicating the status of the video shake correction process, that is, whether or not the video shake correction process is being performed, and if so, what its intensity is. Terminal device 11 displays the status indicated by the status data transmitted from terminal device 12 on its display.
[0032] For example, if user X is watching the video displayed on the screen of terminal device 11 and feels the need to reduce video shake, they check the status of the video shake correction process of terminal device 12 displayed on the screen of terminal device 11. If the video shake correction process is not running in terminal device 12, they operate the control device of terminal device 11 to instruct it to start the video shake correction process. If the video shake correction process is running in terminal device 12, they operate the control device of terminal device 11 to instruct it to increase the intensity of the video shake correction process.
[0033] In response to this operation, terminal device 11 transmits control data to terminal device 12 instructing it to start or increase the intensity of the video shake correction process. Upon receiving the control data transmitted from terminal device 11, terminal device 12 starts the video shake correction process or increases the intensity of the video shake correction process that has already started, in accordance with the control data.
[0034] Subsequently, for example, when the doppelganger Y finishes moving and places the terminal device 12 on a desk or the like, the image displayed on the terminal device 11's screen will no longer be shaken. In this case, the real X operates the control device of the terminal device 11 to instruct it to end the image shake correction process. In response to this operation, the terminal device 11 transmits control data to the terminal device 12 instructing it to end the image shake correction process. When the terminal device 12 receives the control data transmitted from the terminal device 11, it terminates the image shake correction process according to that control data.
[0035] According to the system 1 of the second embodiment, the person X who views the video generated by the camera of the terminal device 12 can directly instruct the terminal device 12 to start and stop the video shake correction process, change the settings of the video shake correction process, and perform other control operations. Therefore, compared to, for example, the case where the person X asks their avatar Y to perform these control operations, it is more convenient for both the person X and the avatar Y, and the person X's judgment is quickly reflected in the video, which is desirable.
[0036] In the above explanation, it was assumed that the instructions given by person X were made by operating the control device, but the instructions given by person X are not limited to operating the control device. For example, person X may give instructions by speaking. In this case, the microphone of terminal device 11 picks up the voice spoken by person X, terminal device 11 recognizes the voice, and transmits control data to terminal device 12 indicating control that follows the recognized instructions of person X.
[0037] Furthermore, in the above explanation, the status regarding the video shake correction processing in terminal device 12 is shown on the display of terminal device 11, but the method of notifying user X of the status is not limited to display. For example, terminal device 11 may synthesize an audio notification of the status indicated by the status data, and user X may be notified of the status by having the speaker of terminal device 11 pronounce the audio.
[0038] [2-1. Modification of the second embodiment] A modified example of System 1 according to the second embodiment is shown below.
[0039] In the second embodiment described above, the status of the video shake correction process in the terminal device 12 is notified to user X from the terminal device 11. The status notified to user X is not limited to the status of the video shake correction process. For example, status data indicating the status of any of the following in the terminal device 12 may be transmitted from the terminal device 12 to the terminal device 11, and the status indicated by that status data may be notified to user X from the terminal device 11.
[0040] Status of the terminal device 12's battery (battery level, etc.) Status of the processor in terminal device 12 (operating clock speed, etc.) Status of the camera on terminal device 12 (pixel count, exposure, etc.) Status of the microphone on terminal device 12 (ON / OFF, sensitivity, directivity, etc.) Status of the display of terminal device 12 (ON / OFF, brightness, etc.) Status of the speaker on terminal device 12 (ON / OFF, volume, etc.) Status of programs running on terminal device 12 (names of running programs, amount of resources on terminal device 12 used by those programs, etc.) Status of the communication device of terminal device 12 (such as ON / OFF status of communication units compliant with communication standards such as WiFi (registered trademark), Bluetooth (registered trademark), and NFC (Near Field Communication)).
[0041] In addition to the status described above, the attributes of terminal device 12 may also be notified to user X from terminal device 11. The attributes of terminal device 12 may include, for example, the model name of terminal device 12, the size of the display, the battery capacity, etc.
[0042] According to this modified version, the user X can easily find out, for example, processes that are wasting the battery of terminal device 12. Furthermore, the user X can, for example, look at the model name of terminal device 12 to find out the battery consumption characteristics of terminal device 12, and instruct their avatar Y in advance to prepare a spare battery to prevent terminal device 12 from unexpectedly running out of battery.
[0043] Furthermore, in the second embodiment described above, the terminal device 12 performs control related to image shake correction processing in accordance with instructions given by user X to the terminal device 11. The objects of control performed by the terminal device 12 in accordance with user X's operations are not limited to image shake correction processing. For example, in accordance with instructions from user X to the terminal device 11, control data instructing control related to any of the following may be transmitted from the terminal device 11 to the terminal device 12, and control in accordance with that control data may be performed in the terminal device 12.
[0044] Control of the processor of terminal device 12 (such as changing the operating clock speed) Control of the camera of terminal device 12 (changing the number of pixels, changing the exposure, etc.) Control of the microphone of terminal device 12 (switching ON / OFF, changing sensitivity, changing directivity, etc.) Control of the display of terminal device 12 (switching ON / OFF, changing brightness, etc.) Control of the speaker of terminal device 12 (switching ON / OFF, changing volume, etc.) Control over programs running on terminal device 12 (termination of running programs, startup of programs not currently running, etc.) Control of the communication device of terminal device 12 (such as switching the operation of communication units ON / OFF in accordance with communication standards such as WiFi (registered trademark), Bluetooth (registered trademark), and NFC (Near Field Communication)).
[0045] According to this modified version, user X can easily terminate processes that are unnecessarily consuming the battery of terminal device 12, or change the settings in terminal device 12 to reduce the battery consumption of terminal device 12.
[0046] [3. Third Embodiment] A third embodiment of the present invention will be described below. The system configuration according to the third embodiment is the same as that of system 1 according to the first embodiment shown in Figure 1. The differences between the third embodiment and the first embodiment will be described below, and the common points will not be explained.
[0047] System 1 according to the third embodiment (hereinafter simply referred to as "System 1") estimates the time during which video shake correction processing is performed in the terminal device 12 during the execution of an activity (hereinafter referred to as "task") that the avatar Y performs at the request of the real person X, based on the scheduled information of the activity (hereinafter referred to as "task") that the avatar Y will perform at the request of the real person X. Based on the estimated time, it determines whether the battery level of the terminal device 12 is sufficient until the completion of the task, and if the battery level of the terminal device 12 is insufficient until the completion of the task, how much capacity of auxiliary battery the avatar Y needs to prepare, and notifies the avatar Y of this determined information.
[0048] Furthermore, some or all of the above processing may be performed by terminal device 11 or by terminal device 12.
[0049] Task scheduling information includes, for example, the following: (Subtask 1) 13:00-13:30 Travel by train from location P1 to location P2 (streaming required) (Subtask 2) 13:30-15:00 Conduct research while moving on foot at location P2 (streaming required) (Subtask 3) 15:00-15:15 Travel by taxi from location P2 to location P3 (no streaming required) (Subtask 4) 15:15~16:00 Interview at a fixed location in P3 (must be streamed)
[0050] As described above, the schedule information includes information for each of the subtasks, which are one or more elements that make up the task, such as the estimated time required, whether or not travel is required, the means of travel if travel is required, and whether or not delivery is necessary.
[0051] For example, based on the above schedule information, System 1 (terminal device 11 or terminal device 12) estimates that for subtask 1, video shake correction processing (correction intensity: weak) is required for 30 minutes, for subtask 2, video shake correction processing (correction intensity: strong) is required for 1 hour and 30 minutes, and for subtasks 3 and 4, video shake correction processing is not required.
[0052] Furthermore, System 1 estimates the operating time of the camera on terminal device 12 during task execution based on the scheduled information. In this case, the camera's operating time is estimated to be 2 hours and 45 minutes.
[0053] System 1 determines whether the battery level of terminal device 12 will be insufficient by the time the task is completed, based on the battery level of terminal device 12 at the start of the task (battery capacity if fully charged), the estimated execution time of the video shake correction process (in the above example, 30 minutes for correction intensity: weak, and 1 hour 30 minutes for correction intensity: strong), the camera's operating time (in the above example, 2 hours 45 minutes), and the total task duration indicated by the schedule information (3 hours in the above example).
[0054] System 1 makes this determination by, for example, calculating an estimated value of the battery level of the terminal device 12 at the time of task completion according to a predetermined calculation formula that includes the battery level at the start of the task, the execution time of the video shake correction process, the camera's operating time, and the total time required for the task (3 hours in the above example) as variables, and comparing this estimated value with a predetermined threshold.
[0055] Alternatively, System 1 may perform this determination using artificial intelligence. For example, if a machine learning model is used, a machine learning model constructed using the following set of variables obtained as training data from tasks previously performed by the avatar Y using terminal device 12 (tasks requested by the real X, or tasks requested by a real person other than the real X) will be used.
[0056] Explanatory variables: Battery level at task start, execution time of image stabilization process, camera operating time, total task duration. Target variable: Battery level at the end of the task
[0057] System 1 inputs the battery level of terminal device 12 at the start of a task, the estimated execution time of the video shake correction process based on the scheduled information, the camera's operating time, and the total duration of the task into a machine learning model for a task to be executed in the future. If the estimated battery level at the time of task completion output from the machine learning model is above a predetermined threshold, System 1 determines that the battery level of terminal device 12 will not be insufficient until the task is completed. On the other hand, if the estimated battery level at the time of task completion output from the machine learning model is below a predetermined threshold, System 1 determines that the battery level of terminal device 12 will be insufficient until the task is completed.
[0058] System 1 notifies its avatar Y, for example via terminal device 12, of whether or not the battery level will be insufficient to complete the task, as determined above. If avatar Y receives notification from terminal device 12 that the battery level is insufficient, it prepares an auxiliary battery to charge the terminal device 12's battery at location S and brings that auxiliary battery with it when executing the task.
[0059] According to the system 1 of the third embodiment, the avatar Y can easily know in advance whether the battery level of the terminal device 12 is sufficient to complete the task. As a result, the avatar Y can determine whether or not to bring an auxiliary battery when performing the task.
[0060] Furthermore, if System 1 determines that the battery level of terminal device 12 will be insufficient before the task is completed, it may estimate the capacity of the auxiliary battery that avatar Y should bring to location S, and notify avatar Y of the estimated auxiliary battery capacity via terminal device 12.
[0061] For example, the target variable of the above machine learning model is changed from battery level to power consumption. System 1 then estimates the minimum battery level that the auxiliary battery that avatar Y should bring to location S must meet by subtracting the battery level of terminal device 12 at the start of the task from the estimated power consumption output by the machine learning model that terminal device 12 will need until the task is completed (or by adding a predetermined amount to that value). Terminal device 12 then notifies avatar Y of the estimated battery level of the auxiliary battery.
[0062] In this case, the clone Y can easily find out the remaining charge of the auxiliary battery that should be brought to location S.
[0063] [3-1. Modified Examples of the Third Embodiment] A modified example of System 1 according to the third embodiment is shown below.
[0064] In the third embodiment described above, the information used to determine whether the battery level of the terminal device 12 will be insufficient by the time the task is completed is the battery level of the terminal device 12 at the start of the task, the execution time of the video shake correction process, the operating time of the camera, and the total time required for the task. However, in addition to or instead of this information, other information may be used to determine whether the battery level of the terminal device 12 will be insufficient by the time the task is completed. For example, the information used to determine whether the battery level of the terminal device 12 will be insufficient by the time the task is completed may include the operating time of the microphone and the operating time of the speaker.
[0065] Furthermore, during task execution, the system may repeatedly determine, for example, at sufficiently short time intervals, whether the battery level of terminal device 12 will be insufficient by the time the task is completed, and notify the avatar Y of the result. In this case, the system 1 uses the battery level of terminal device 12 at the time the determination is made to make the determination. Alternatively, the system 1 may predict future changes in the battery level of terminal device 12 over time based on past changes in the battery level of terminal device 12 during task execution, and based on that prediction, determine whether the battery level will be insufficient by the time the task is completed.
[0066] Furthermore, the result of the determination of whether or not the battery level is low may be notified to the original user X by the terminal device 11, in addition to or instead of to the avatar Y. In this case, if it is predicted that the battery level of the terminal device 12 will not last until the completion of the task, the original user X can take some measures to prevent the task from being interrupted due to insufficient battery power.
[0067] [4. Fourth Embodiment] A fourth embodiment of the present invention will be described below. The system configuration according to the fourth embodiment is the same as that of system 1 according to the first embodiment shown in Figure 1. The differences between the fourth embodiment and the first embodiment will be described below, and the common points will not be explained.
[0068] System 1 according to the fourth embodiment (hereinafter simply referred to as "System 1") automatically controls the video shake correction processing performed in the terminal device 12, that is, without requiring any operation from the person X or their avatar Y.
[0069] In System 1, for example, task schedule information used in the third embodiment is used. As previously described, task schedule information is, for example, the following information. (Subtask 1) 13:00-13:30 Travel by train from location P1 to location P2 (streaming required) (Subtask 2) 13:30-15:00 Conduct research while moving on foot at location P2 (streaming required) (Subtask 3) 15:00-15:15 Travel by taxi from location P2 to location P3 (no streaming required) (Subtask 4) 15:15~16:00 Interview at a fixed location in P3 (must be streamed)
[0070] The terminal device 12 starts, ends, or adjusts the intensity of the image stabilization process according to the schedule information. For example, if the terminal device 12 follows the above schedule information, it will start the image stabilization process with a weak intensity at 13:00 on the task execution day, then change the intensity of the ongoing image stabilization process to strong at 13:30, and then end the image stabilization process at 15:00.
[0071] According to the system 1 of the fourth embodiment, the clone Y does not need to perform operations on the terminal device 12 such as starting, ending, or changing the correction intensity of the video shake correction process.
[0072] In the above description, control related to the video shake correction process is assumed to be performed by the terminal device 12 according to the task schedule information. However, as in the second embodiment, the terminal device 11 may generate control data instructing the control related to the video shake correction process at the appropriate timing according to the task schedule information and transmit it to the terminal device 12. In this case, the terminal device 12 performs actions such as starting, ending, and adjusting the correction intensity of the video shake correction process according to the control data transmitted from the terminal device 11.
[0073] [4-1. Modified Examples of the Fourth Embodiment] A modified example of System 1 according to the fourth embodiment is shown below.
[0074] In the fourth embodiment described above, control of the image shake correction process is performed automatically, that is, without requiring any operation from the person X or their avatar Y. Alternatively, control of the image shake correction process may be performed semi-automatically, that is, in response to simple operations by the person X or their avatar Y.
[0075] Tasks do not always proceed as planned according to the schedule information. Therefore, in this modified example, terminal device 12 displays a button on its display for the avatar Y to instruct the execution of the control when it is time for the control indicated by the schedule information to be performed (or a predetermined time period earlier than that time). For example, five minutes before 13:30, the scheduled start time of subtask 2, terminal device 12 displays a virtual operation button ("Strong" button) on its display for changing the correction intensity of the ongoing video shake correction process from weak to strong. Then, at the time when subtask 2 actually starts, avatar Y can appropriately adjust the correction intensity of the video shake correction process by operating the displayed "Strong" button.
[0076] Furthermore, in the fourth embodiment described above, control related to the image shake correction process is performed automatically or semi-automatically, triggered by the arrival of the timing indicated by the task schedule information. In addition to this, or instead, control related to the image shake correction process may be performed automatically or semi-automatically, triggered by the occurrence of other types of events.
[0077] For example, if the ambient light around the camera of the terminal device 12 is insufficient, it is difficult to eliminate image shake using electronic image shake correction processing. Therefore, if the terminal device 12 is equipped with an illuminance sensor to measure ambient light, the terminal device 12 may automatically start image shake correction processing when the ambient light measured by the illuminance sensor exceeds a predetermined threshold, and automatically end the image shake correction processing when the ambient light measured by the illuminance sensor falls below the predetermined threshold.
[0078] Furthermore, if the camera of the terminal device 12 is not shaking, image shake correction processing is unnecessary. Therefore, if the terminal device 12 is equipped with a motion sensor that detects shaking of the device (or camera), the terminal device 12 may automatically start image shake correction processing when the motion sensor detects shaking of the device, and automatically end the image shake correction processing when a predetermined period of time has passed during which the motion sensor does not detect shaking of the device.
[0079] Furthermore, in the fourth embodiment described above, the object that is automatically (or semi-automatically) controlled is the image shake correction process, but other processes besides the image shake correction process may also be automatically (or semi-automatically) controlled.
[0080] For example, as shown in Figure 2, the avatar Y may show the display of terminal device 12 to person Z at location S, and communication may take place between the real person X and person Z via terminal devices 11 and 12. In addition to the avatar Y and person Z, there may be other people unrelated to the real person X, such as passersby, at location S, and these people may be having conversations. In such cases, it is desirable for the real person X if the microphone of terminal device 12 is operational and has high sensitivity while person Z is speaking, and deactivated or has low sensitivity when person Z is not speaking, so that the voices of unrelated people are not audible at high volume from the speaker of terminal device 11.
[0081] Therefore, the terminal device 12 determines whether or not person Z is speaking based on the movement of person Z's mouth (even through a mask) as seen in the image generated by the camera. During the period when it is determined that person Z is speaking, the sensitivity of the microphone of the terminal device 12 is automatically increased, and during the period when it is determined that person Z is not speaking, the sensitivity of the microphone of the terminal device 12 is automatically decreased.
[0082] Alternatively, instead of increasing or decreasing the microphone sensitivity, the terminal device 12 may turn the microphone ON or OFF.
[0083] Furthermore, the terminal device 12 may detect, based on the image generated by the camera, that a person Z in the image is listening intently to the sound coming from the terminal device 12, and automatically increase the volume of the speaker of the terminal device 12 as a trigger. The action of listening intently to the sound coming from the terminal device 12 includes, for example, bringing one's ear close to the speaker of the terminal device 12 or placing one's hand behind one's ear.
[0084] Furthermore, if the terminal device 12 is equipped with multiple speakers that emit sounds at different volumes, the terminal device 12 may, upon detecting that person Z in the video is listening intently to the sound coming from the terminal device 12, switch the speaker emitting the sound from the speaker with the lower volume to the speaker with the higher volume.
[0085] Furthermore, terminal device 12 may estimate the emotions of person Z at that time from the facial expressions and gestures of person Z captured in the image generated by the camera, and send the estimated emotion, or a message corresponding to that emotion, to terminal device 11. In this case, the message might be something like, "The other person seems a little unhappy." Terminal device 11 displays the emotion of person Z, or a message corresponding to that emotion, transmitted from terminal device 12, on its display. Alternatively, terminal device 11 synthesizes audio representing the emotion of person Z, or a message corresponding to that emotion, transmitted from terminal device 12, and pronounces that audio through its speaker. Thus, person X can easily become aware of person Z's emotions.
[0086] Furthermore, the terminal device 12 may estimate the emotions of person Z at that time based on the voice of person Z picked up by the microphone, in addition to or instead of the video generated by the camera, and send the estimated emotions, or a message corresponding to those emotions, to the terminal device 11.
[0087] Furthermore, the terminal device 12 may estimate the situation at location S based on the video generated by the camera or the sound picked up by the microphone, and send the estimated situation, or a message corresponding to that situation, to the terminal device 11. In this case, the message might be something like, "It's very quiet around here." In this case, person X can easily learn about the situation at location S.
[0088] Furthermore, the terminal device 12 may increase or decrease the microphone sensitivity, or turn the microphone ON or OFF, based on the sound picked up by the microphone.
[0089] For example, if person X speaks continuously for a predetermined period of time or longer, and person Z does not speak during that time, terminal device 12 will determine that person X is speaking unilaterally to person Z and will maintain a low microphone sensitivity. Also, if terminal device 12 determines that person X has asked a question, such as "What product do you recommend?", it will predict that person Z will speak next and increase the microphone sensitivity of terminal device 12. As for how terminal device 12 determines whether the words spoken by person X are a question, it may employ any of the following methods: speech recognition of the content of person X's speech and determining whether the content is a question; or determining whether the intonation of person X's speech is the intonation of a question (for example, an intonation that rises at the end of a sentence).
[0090] Furthermore, if System 1 (terminal device 11 or terminal device 12) recognizes the content of Person X's speech and determines that the recognized speech is a question directed to Person Z, it may display an answer input form corresponding to that question on the display of Terminal Device 12. For example, if Person X asks Person Z a question such as, "How much does this product cost?", the display of Terminal Device 12 will display an answer input form that includes, for example, the question "How much does this product cost?", a text box for entering a price such as "yen", and a virtual numeric keypad for Person Z to input numbers. Person Z can easily answer Person X by entering their answer into this answer input form.
[0091] Similarly, when person Z asks person X a question, an answer input form corresponding to that question may be displayed on the terminal device 11's display.
[0092] Furthermore, if person X and person Z use different languages, it is desirable that the language used in the response input form be translated into the language used by the responding user. Similarly, it is desirable that the content of the response notified to the user asking the question be translated into the language used by the user asking the question. In this case, even if person X and person Z use different languages, questions and answers can be asked smoothly.
[0093] Furthermore, the type of object that accepts responses in the response input form is not limited to text boxes. For example, a menu list that prompts the user to select one or more options from multiple choices may also be included in the response input form.
[0094] Furthermore, if person X utters specific words such as person Z's name, system 1 (terminal device 11 or terminal device 12) may detect this and turn on the microphone of terminal device 12 or increase the sensitivity of a microphone that is already on. At the same time as turning on the microphone or increasing its sensitivity, or after a predetermined period of time has elapsed, terminal device 12 may also turn on the camera of terminal device 12, which was previously off.
[0095] Furthermore, before the terminal device 12 performs the above-described control on devices such as the microphone and camera, it may display a confirmation message on the display, for example, "Turning on the camera. Is that OK? Yes / No," and the terminal device 12 may only execute the control if person Z performs an action indicating their understanding of the confirmation message (for example, by touching the "Yes" button).
[0096] As described above, in communication using System 1, the user can include an identifier such as the name of the person they want to draw attention to in their utterance, thereby causing the participant's terminal device identified by that identifier to perform an action corresponding to the content of the utterance.
[0097] For example, as shown in Figure 3, if person V is also present at location S in addition to person Z, and person V is participating in the communication between person X and person Z using terminal device 13, person X can, for example, say "Z, where is your hometown?" to display an input form on the terminal device 12 prompting the user to enter the name of their hometown, and for example, say "V, how many people are in your family?" to display an input form on the terminal device 13 prompting the user to enter the number of family members.
[0098] Furthermore, the terminal device 12 may operate in a registration mode, which is an operating mode for registering a speaker in response to a predetermined operation by, for example, the avatar Y, and the system 1 (terminal device 11 or terminal device 12) may be able to identify the participants in the communication based on the sound picked up by the microphone of the terminal device 12 while it is operating in registration mode.
[0099] As shown in Figure 4, there are doppelgangers Y, person Z, person V, and person W at location S, and the real person X communicates with person Z and person V. Person W is a passerby unrelated to the real person X, etc. Person Z and person V communicate with the real person X via terminal device 12.
[0100] In this case, Person Z and Person V give brief self-introductions to the microphone of the terminal device 12, which is operating in registration mode, for example, in that order. Alternatively, Person Z and Person V may read a pre-written sentence instead of introducing themselves. In that case, it is desirable that the pre-written sentence include various types of pronunciation so that the characteristics of the speakers' voices can be easily grasped. System 1 extracts features from the voice picked up by the microphone of the terminal device 12 during registration mode and stores the extracted features. After that, the avatar Y operates the terminal device 12 to switch the operating mode of the terminal device 12 to normal mode.
[0101] Subsequently, the terminal device 12 increases the microphone sensitivity during periods when the sound picked up by the terminal device 12's microphone includes the voice of person Z or person V, and decreases the microphone sensitivity during other periods (for example, when the voice of person W is included, but the voice of person Z or person V is not). This allows person X to concentrate on the conversation with the person they are communicating with at location S, even if the location S is noisy due to other passersby, etc.
[0102] [5. Categories of Inventions] The present invention provides a system exemplified by System 1 in the above-described embodiment, a terminal device exemplified by Terminal Device 11 or Terminal Device 12, a method executed by System 1, a program for causing a computer constituting Terminal Device 11 or Terminal Device 12 to perform data processing performed by Terminal Device 11 or Terminal Device 12, a recording medium on which the program is recorded, or a data processing device that continuously stores the program. [Explanation of symbols]
[0103] 1...System, 9...Communication network, 11...Terminal device, 12...Terminal device, 13...Terminal device.
Claims
[Claim 1] A system comprising a first terminal device used by the person concerned and a second terminal device used by a doppelganger, which is a person or device present at the location acting as a substitute for the person concerned, includes the step of the second terminal device streaming video data representing images generated by a camera carried by the doppelganger to the first terminal device via a communication network, The system performs image shake correction processing, which is data processing to reduce image shake contained in the image represented by the image shake correction processing, on the image data streamed from the first terminal device via the communication network, and displays the image represented by the image shake correction processing on a display. A method for providing this.
Citation Information
Patent Citations
Remote imaging system
JP2008219310A
Telepresence system, flight vehicle control program, and movable traveling body control program
JP2021096576A
Experience sharing system, experience sharing method
JP6644288B1
Camerawork generating method and video processing device
WO2018030206A1