Terminal device
The terminal device improves the realism of virtual face-to-face calls by generating and transmitting detailed model images of users, addressing the limitations of existing technologies in multi-user scenarios.
Patent Information
- Application Number
- JP2022212662
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2042-12-28
AI Technical Summary
Existing virtual face-to-face call technologies lack the realism needed for immersive communication, particularly in multi-user scenarios where the representation of users is not detailed enough.
A terminal device equipped with a display unit, multiple imaging units, and a control unit that generates and transmits model images representing users based on captured images, allowing for more detailed and realistic user representation in virtual space.
Enhances the realism of virtual face-to-face calls by providing detailed and dynamic model images of users, improving the overall immersive experience even in large-scale virtual meetings.
Smart Images

Figure 0007694555000001 
Figure 0007694555000002 
Figure 0007694555000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a terminal device.
Background Art
[0002] A method is known in which users at different locations communicate via a network using a computer, and perform a virtual face-to-face call by transmitting and receiving each other's images and voices. Various techniques for supporting face-to-face calls on such a network have been proposed. For example, Patent Document 1 discloses a technique for selecting a captured image of a caller using information regarding the position and orientation of the caller with respect to the screen.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] There is room for further improving the reality of virtual face-to-face calls.
[0005] Hereinafter, a terminal device and the like that enable improvement of the reality of virtual face-to-face calls will be disclosed.
Means for Solving the Problems
[0006] The terminal device in the present disclosure is a terminal device including a display unit capable of displaying an image toward a user in front, a first imaging unit provided around the display unit, a plurality of second imaging units provided behind the display unit, a communication unit, and a control unit that performs communication through the communication unit. The control unit sends information for generating a model image representing the user based on the captured image of the second imaging unit corresponding to the position of the user included in the captured image of the first imaging unit to another terminal device so that the other terminal device can display the model image.
Advantages of the Invention
[0007] According to the terminal device and the like in the present disclosure, it is possible to improve the reality in virtual face-to-face calls.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Modes for Carrying Out the Invention
[0009] Hereinafter, embodiments will be described.
[0010] With reference to FIGS. 1 and 2, a configuration example of the call system 1 in an embodiment will be described. FIG. 1 shows a configuration example of the entire call system, and FIG. 2 shows a configuration example of the terminal device 12 included in the call system 1.
[0011] The call system 1 includes a server device 10 and a plurality of terminal devices 12 that are connected to each other via a network 11 so as to be able to communicate information. The call system 1 is a system for providing a call event in a virtual space in which a user can participate using the terminal device 12. In the call event in the virtual space, each user is represented by a model image representing each of them.
[0012] The server device 10 belongs to, for example, a cloud computing system or other computing systems, and is a server computer that functions as a server for implementing various functions. The server device 10 may be composed of two or more server computers that are communicably connected and operate in cooperation. The server device 10 executes transmission and reception of information necessary for providing a call event and information processing.
[0013] The terminal device 12 is a device capable of information processing and information communication, and has an image display and input function, an imaging function, and a voice input / output communication function, and is used by a user who participates in a call in a virtual space provided by the server device 10. The terminal device 12 includes, for example, a display device including a transmissive touch display, a camera, a microphone, and a speaker, and an information processing device such as a personal computer having a communication interface.
[0014] As shown in FIG. 2 for example, the terminal device 12 includes a transmissive display 22 capable of displaying an image of a call partner facing users 25 and 26, and a transmissive touch panel 23 integrally provided so as to overlap the display 22 and enabling touch input. The display 22 is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display. The display 22 and the touch panel 23 have a size such as, for example, about 100 centimeters in height and 150 centimeters in width, which enables multiple users 25 and 26 to line up and simultaneously view and draw images. This size is an example and is not limited to the numerical values shown here. Further, the terminal device 12 includes an overhead camera 20 provided around the display 22, for example, at the upper part, and imaging the front area of the display 22 where the users 25 and 26 are located. Further, the terminal device 12 is provided at a position, for example, from a dozen centimeters to several tens of centimeters behind the display 22 and the touch panel 23, and has two or more arbitrary numbers of in-screen cameras 21 capable of imaging the front area of the display 22 through the display 22 and the touch panel 23. The plurality of in-screen cameras 21 are installed, for example, at a height capable of imaging the upper bodies of the users 25 and 26 at arbitrary intervals in the horizontal direction, for example, at a height of 100 centimeters to 150 centimeters from the floor or the ground.
[0015] The network 11 is, for example, the Internet, but includes an ad-hoc network, a LAN (Local Area Network), a MAN (Metropolitan Area Network), or other networks or any combination thereof.
[0016] In this embodiment, the terminal device 12 sends information for generating model images representing the users 25 and 26 to other terminal devices 12 so that the other terminal devices 12 can display the respective model images, based on the captured images of the in-screen cameras 21 corresponding to the positions of the users 25 and 26 included in the captured image of the overhead camera 20. The model images are switched between 2D models and 3D models according to the distances of the users 25 and 26 from the in-screen cameras 21. Therefore, even in a case where a relatively large display 22 is used simultaneously by a plurality of users 25 and 26, by identifying the in-screen camera 21 closest to the users 25 and 26 based on the captured image by the overhead camera 20 and using the captured image thereof, it becomes possible to obtain information for generating image models that represent each user in more detail. By transmitting and receiving such information between the terminal devices 12, it becomes possible to improve the reality of face-to-face calls in the virtual space.
[0017] The configurations of the server device 10 and the terminal device 12 will be described in detail below.
[0018] The server device 10 includes a communication unit 101, a storage unit 102, and a control unit 103. These configurations are appropriately arranged in two or more computers when the server device 10 is composed of two or more server computers.
[0019] The communication unit 101 includes one or more communication interfaces. The communication interface is, for example, a LAN interface. The communication unit 101 receives information used for the operation of the server device 10 and transmits information obtained by the operation of the server device 10. The server device 10 is connected to the network 11 by the communication unit 101 and communicates with the terminal device 12 via the network 11.
[0020] The storage unit 102 includes, for example, one or more semiconductor memories that function as a main memory device, an auxiliary memory device, or a cache memory, one or more magnetic memories, one or more optical memories, or a combination of at least two of these. The semiconductor memory is, for example, a RAM (Random Access Memory) or a ROM (Read Only Memory). The RAM is, for example, an SRAM (Static RAM) or a DRAM (Dynamic RAM). The ROM is, for example, an EEPROM (Electrically Erasable Programmable ROM). The storage unit 102 stores information used for the operation of the control unit 103 and information obtained by the operation of the control unit 103.
[0021] The control unit 103 includes one or more processors, one or more dedicated circuits, or a combination of these. The processor is, for example, a general-purpose processor such as a CPU (Central Processing Unit), or a dedicated processor such as a GPU (Graphics Processing Unit) specialized for specific processing. The dedicated circuit is, for example, an FPGA (Field-Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), etc. The control unit 103 executes information processing related to the operation of the server device 10 while controlling each part of the server device 10.
[0022] The functions of the server device 10 are realized by a processor included in the control unit 103 executing a control program. The control program is a program for causing a computer to function as the server device 10. Also, some or all of the functions of the server device 10 may be realized by a dedicated circuit included in the control unit 103. Further, the control program may be stored in a non-transitory recording and storage medium readable by the server device 10, and the server device 10 may read it from the medium.
[0023] The terminal device 12 includes a communication unit 111, a memory unit 112, a control unit 113, an input unit 115, a display / output unit 116, and an imaging unit 117. The input unit 115 includes a touch panel 23. The display / output unit 116 includes a display 22. The imaging unit 117 includes an overhead camera 20 and an in-screen camera 21.
[0024] The communication unit 111 has a communication module corresponding to a wired or wireless LAN standard, a module corresponding to a mobile communication standard such as LTE, 4G, 5G, etc. The terminal device 12 is connected to the network 11 via the communication unit 111 through a nearby router device or a base station for mobile communication, and performs information communication with the server device 10 etc. via the network 11.
[0025] The memory unit 112 includes one or more semiconductor memories, one or more magnetic memories, one or more optical memories, or a combination of at least two of these. The semiconductor memory is, for example, a RAM or a ROM. The RAM is, for example, an SRAM or a DRAM. The ROM is, for example, an EEPROM. The memory unit 112 functions as, for example, a main memory device, an auxiliary memory device, or a cache memory. The memory unit 112 stores information used for the operation of the control unit 113 and information obtained by the operation of the control unit 113.
[0026] The control unit 113 has, for example, one or more general-purpose processors such as a CPU or an MPU (Micro Processing Unit), or one or more dedicated processors such as a GPU specialized for specific processing. Alternatively, the control unit 113 may have one or more dedicated circuits such as an FPGA or an ASIC. The control unit 113 comprehensively controls the operation of the terminal device 12 by operating according to a control / processing program or operating according to an operation procedure implemented as a circuit. Then, the control unit 113 transmits and receives various information to and from the server device 10 etc. via the communication unit 111, and executes the operation according to this embodiment.
[0027] The input unit 115 includes one or more input interfaces. The input interface is, for example, a touch panel 23 superimposed on or integrally provided with the display 22. The input interface also includes physical keys, capacitive keys, a pointing device, and a microphone for receiving voice input. Further, the input interface may include a scanner or a camera for scanning an image code, and an IC card reader. The input unit 115 receives an operation for inputting information used for the operation of the control unit 113, and sends the input information to the control unit 113.
[0028] The display / output unit 116 includes one or more output interfaces. The output interface is, for example, the display 22. The output interface also includes a speaker. The display / output unit 116 outputs information obtained by the operation of the control unit 113.
[0029] The imaging unit 117 includes an overhead camera 20 and a plurality of in-screen cameras 21. The overhead camera 20 and each in-screen camera 21 each include a visible light camera that captures a captured image of a subject using visible light, and a distance measurement sensor that measures the distance to the subject and acquires a distance image. The visible light camera captures the subject, for example, at 15 to 30 frames per second to generate a moving image composed of continuous captured images. The distance measurement sensor includes a ToF (Time Of Flight) camera, LiDAR (Light Detection And Ranging), and a stereo camera, and generates a distance image of the subject including distance information. The overhead camera 20 sends the captured image (hereinafter referred to as the overhead image) and the distance image (hereinafter referred to as the overhead distance image) to the control unit 113. Each in-screen camera 21 also sends the captured image (hereinafter referred to as the approach image) and the distance image (hereinafter referred to as the approach distance image) to the control unit 113.
[0030] The functions of the control unit 113 are realized by a processor included in the control unit 113 executing a control program. The control program is a program for causing the processor to function as the control unit 113. Also, some or all of the functions of the control unit 113 may be realized by a dedicated circuit included in the control unit 113. Further, the control program may be stored in a non-transitory recording and storage medium readable by the terminal device 12, and the terminal device 12 may read it from the medium.
[0031] FIGS. 3A and 3B are flowchart diagrams for explaining the operation procedure of the terminal device 12 related to the execution of a call event. The procedures shown here are executed by the control unit 113 when a plurality of terminal devices 12 establish a connection with each other via the server device 10 and then transmit and receive information and voice information for displaying model images of their respective users.
[0032] FIG. 3A relates to the operation procedure of the control unit 113 when each terminal device 12 sends out information and the like for generating a model image of the user of that terminal device 12.
[0033] In step S300, the control unit 113 causes the overhead camera 20 of the imaging unit 117 to capture an overhead image of the user and acquire an overhead distance image at an arbitrarily set frame rate, and causes the input unit 115 to collect the voice of the user's speech. The control unit 113 acquires the overhead image and the distance image from the imaging unit 117, and acquires voice information from the input unit 115.
[0034] In step S301, the control unit 113 derives the position of each user based on the bird's-eye view image and the bird's-eye view distance image. The control unit 113 detects the user by image processing such as pattern recognition from the bird's-eye view image. Further, the control unit 113 derives the distance and direction from the bird's-eye view camera 20 at, for example, the position of the center of the face of each detected user. The center of the face may be the center of both eyes further detected within the face, or may be the center or centroid of a geometric region detected as the face. Then, the control unit 113 derives the spatial coordinates of each user using the spatial coordinates of the bird's-eye view camera 20 stored in the storage unit 112 in advance.
[0035] In step S302, the control unit 113 causes the in-screen camera 21 closest to each user to capture an approach image and acquire an approach distance image respectively. The control unit 113 derives the in-screen camera 21 closest to each user using the spatial coordinates of each in-screen camera 21 stored in the storage unit 112 in advance and the spatial coordinates of each user. The control unit 113 selects the in-screen camera 21 closest to each user by, for example, deriving the distance of the combination of the center of the face of each user and the position of the lens of each in-screen camera 21 and searching for the shortest one. Then, the control unit 113 activates the selected in-screen camera 21 and causes it to capture an approach image and acquire an approach distance image at an arbitrarily set frame rate. By stopping the operation of the in-screen camera 21 that is not selected, it is possible to reduce power consumption.
[0036] In step S303, the control unit 113 determines the type of the model image representing each user. The control unit 113 determines the type of the model image as a 2D model or a 3D model according to the distance between each user and the in-screen camera 21 closest to the user. The control unit 113 uses an arbitrary criterion stored in the storage unit 112 in advance and determines that it is a 2D model when the distance between each user and the in-screen camera 21 closest to the user is less than the criterion, and a 3D model when it is greater than or equal to the criterion. The criterion is, for example, several tens of centimeters to one hundred and several tens of centimeters.
[0037] In step S304, the control unit 113 encodes the proximity image, the proximity distance image, the model information, and the voice information to generate encoded information. The model information is the type of the model image of each user and the information indicating the arrangement of each user. The information indicating the arrangement of each user is, for example, the information on the position of the in-screen camera 21 that captured the proximity image. The control unit 113 may perform arbitrary processing (such as resolution change and trimming) on the captured image or the like during encoding.
[0038] In step S305, the control unit 113 packetizes the encoded information by the communication unit 111 and sends it to the server device 10 for another terminal device 12.
[0039] When the control unit 113 acquires information input in response to an operation by the own user for interrupting imaging and sound collection or exiting a call event (Yes in S306), it ends the processing procedure in FIG. 3A. While the control unit 113 does not acquire information corresponding to an operation for interruption or exit (No in S306), it executes steps S300 to S305 and sends information for generating a model image representing the own user and information for outputting voice to another terminal device 12.
[0040] FIG. 3B relates to the operation procedure of the control unit 113 when the terminal device 12 outputs the image and voice of another user. When the control unit 113 receives a packet sent by another terminal device 12 executing the procedure in FIG. 3A via the server device 10, it executes steps S310 to S316.
[0041] In step S310, the control unit 113 decodes the encoded information included in the packet received from another terminal device 12 to acquire the proximity image, the proximity distance image, the model information, and the voice information.
[0042] In step S312, the control unit 113 generates a model image representing another user. The control unit 113 determines whether to generate a 2D model or a 3D model based on the model information. Then, the control unit 113 generates a 2D model or a 3D model representing each user based on the proximity image, or based on the proximity image and the proximity distance image. When generating the 3D model, the control unit 113 generates a polygon model using the distance image of the other user, and generates a 3D model of the other user by applying texture mapping using the captured image of the other user to the polygon model. However, the generation of the 3D model is not limited to the example shown here, and any method can be adopted.
[0043] In step S313, the control unit 113 arranges models representing each user in the virtual space where the call event is held. The control unit 113 arranges the generated model of the other user at the coordinates in the virtual space. The control unit 113 arranges each model at the corresponding coordinates in the virtual space using the information indicating the arrangement of the models included in the model information. For example, each model is arranged at the position of the in-screen camera 21 corresponding to each model in the display image.
[0044] In step S314, the control unit 113 generates a virtual space image obtained by capturing the models for each user arranged in the virtual space from a virtual viewpoint by rendering.
[0045] In step S316, the control unit 113 displays a display image and outputs sound through the display / output unit 116. That is, the control unit 113 outputs information for displaying an image of the virtual space in which the 2D model or 3D model is arranged in the virtual space to the display / output unit 116. The display / output unit 116 displays the image of the virtual space on the display 22 and outputs sound.
[0046] By repeatedly executing steps S310 to S316 by the control unit 113, the own user can listen to the voice of the other user while watching a video of the virtual space image including the 2D model or 3D model of the other user.
[0047] When the user is relatively close to the in-screen camera 21, the terminal device 12 can reduce the processing load of the control unit 113 by displaying a 2D model that captures the front part of the user. On the other hand, when the user moves away from the in-screen camera 21, the terminal device 12 can improve the reality of the virtual face-to-face call by displaying a 3D model with depth and three-dimensionality.
[0048] FIG. 4 shows the configuration of the terminal device 12 in a modified example. In the modified example, the terminal device 12 has a projector 40 and a transmissive screen 22' that receives the projection light from the projector 40 instead of the transmissive display 22. The projector 40 is installed behind the screen 22' with respect to the user 43 in front of the screen 22'. In such a configuration, the in-screen camera 21 or the peripheral structure such as the wiring of the in-screen camera may interfere with the projection light from the projector 40, causing a shadow in the image 41 projected on the screen 22' and giving the user 43 a sense of discomfort. Therefore, the control unit 113 forms a mask image 42 at the position corresponding to the in-screen camera 21 in the projected image 41'. The mask image 42 is formed to have a lower luminance than the surrounding images and a smaller difference from the shadow, so that the sense of discomfort can be reduced. Further, the control unit 113 generates the mask image 42 in a number corresponding to the number of in-screen cameras 21, each corresponding to the respective in-screen camera 21. Furthermore, the control unit 113 may store the position where the mask image 42 is generated in the storage unit 112 and perform control to accept a drawing operation to a location other than the mask image 42 and not accept a drawing operation to the position of the mask image 42.
[0049] In a further modified example, the terminal device 12 may use an opaque screen such as a whiteboard instead of the transmissive screen 22', and may be configured to provide a projector behind the user 43 and project an image onto the screen. In that case, a hole may be provided at the position corresponding to the in-screen camera 21 of the screen, and the user 43 may be imaged through the hole from the in-screen camera 21.
[0050] The step of generating a model image based on a proximity image may be appropriately distributed to the procedures of FIGS. 3A and 3B. For example, between steps S303 and S304 of FIG. 3A, the control unit 113 may insert a step of generating a model image of the user himself / herself, and encode and transmit the data of the generated model image together with the information on the arrangement of the model and the voice information. When the control unit 113 receives such encoded information from another terminal device 12, it can decode the data of the model image and arrange it in the virtual space according to the arrangement information.
[0051] In the above-described embodiment, the processing / control program that defines the operation of the control unit 113 of the terminal device 12 may be stored in the storage unit of the server device 10 or another server device and downloaded to each terminal device 12 via the network 11, or may be stored in a recordable / memorable medium readable by each terminal device 12, and each terminal device 12 may read it from the medium.
[0052] In the above, the embodiments have been described based on the drawings and examples. It should be noted that those skilled in the art can easily make various modifications and corrections based on the present disclosure. Therefore, it should be noted that these modifications and corrections are included in the scope of the present disclosure. For example, the functions included in each means, each step, etc. can be rearranged so as not to be logically contradictory, and it is possible to combine or divide a plurality of means, steps, etc. into one.
Description of Reference Numerals
[0053] 1 Call system 10 Server device 11 Network 12 Terminal device 101, 111 Communication unit 102, 112 Storage unit 103, 113 Control unit 115 Input unit 116 Display / output unit 117 Imaging unit 20 Overhead camera 21 In-screen camera 22 Display 23 Touch panel
Claims
1. A display unit capable of displaying an image facing a user in front; A first imaging unit provided around the display unit; A plurality of second imaging units provided behind the display unit; A communication unit; A terminal device having a control unit that communicates via the communication unit, The control unit sends information for generating a model image representing the user based on the captured image of the second imaging unit corresponding to the position of the user included in the captured image of the first imaging unit to another terminal device for the other terminal device to display the model image, The model image representing the user is a 2D model when the user is located at a first distance from the display unit and a 3D model when the user is located at a second distance greater than the first distance. Terminal device.
2. In Claim 1, The display unit is a screen that displays a projection image projected from a projector located behind the plurality of second imaging units, The control unit causes the projector to project a mask image at a position where the shadows of the plurality of second imaging units in the projection image interfere, Terminal device.
3. In Claim 2, The control unit makes the luminance of the mask image smaller than the luminance of the image around the mask image in the projection image, Terminal device.
4. In Claim 1, The control unit detects the face of the user included in the captured image of the first imaging unit as the position of the user, causes the second imaging unit corresponding to the position of the user to capture an image of the user, and stops the imaging of the other second imaging units. Terminal device.
Citation Information
Patent Citations
Image communication device
JP2017022600A
3D Telepresence System
JP2019533324A
Immersive teleconferencing with translucent video stream
US20170019627A1
Information processing device, information processing system, information processing method, and program
WO2017195513A1