Terminal device

The terminal device improves stereoscopic image quality in virtual or augmented reality by using its control unit to avoid configuring observation images in overlapping fields of view, thus preventing image mixing and maintaining clarity for multiple users.

JP7687330B2Active Publication Date: 2025-06-03TOYOTA JIDOSHA KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022212682
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-12-28
Publication Date
2025-06-03
Estimated Expiration
2042-12-28

AI Technical Summary

Technical Problem

Existing light field displays for virtual or augmented reality struggle to maintain high image quality when displaying stereoscopic images to multiple users, as the overlap of fields of view can lead to mixed observation images and reduced clarity.

Method used

A terminal device equipped with an imaging unit, a display unit, and a control unit that captures and displays three-dimensional objects for each user's viewpoint position, while avoiding the configuration of observation images in regions where the fields of view overlap, thereby preventing image mixing and maintaining clarity.

Benefits of technology

This solution effectively improves the quality of stereoscopic images by preventing the mixing of observation images in overlapping fields of view, thereby enhancing the user experience in virtual or augmented reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007687330000001
    Figure 0007687330000001
  • Figure 0007687330000002
    Figure 0007687330000002
  • Figure 0007687330000003
    Figure 0007687330000003
Patent Text Reader

Abstract

To provide a terminal capable of improving the quality of stereoscopic images.SOLUTION: The terminal includes: an imaging unit for picking up an image of multiple users; a display unit that displays a three-dimensional object constituted of rays to project an observed image according to the viewpoint position of a three-dimensional object toward the viewpoint position; and a control unit that is configured so as to, when causing the display unit to display different three-dimensional objects to each user's viewpoint position obtained from the captured images of the multiple users no three-dimensional object is formed in the area where the views from each user's viewpoint overlap.SELECTED DRAWING: Figure 2B
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a terminal device.

Background Art

[0002] As an example of a technology for assisting a user experience in virtual reality or augmented reality, various technologies for displaying a stereoscopic image of various three-dimensional objects have been proposed to improve the reality of the user experience. For example, Patent Document 1 discloses a technology related to a light field display that displays a three-dimensional object to a plurality of users.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When displaying a stereoscopic image to a plurality of users by a light field display, there is room for improving the image quality.

[0005] The present disclosure provides a terminal device and the like that enables improvement of the quality of a stereoscopic image.

Means for Solving the Problems

[0006] The terminal device in the present disclosure includes an imaging unit that images a plurality of users, a display unit that displays the three-dimensional object by configuring it with light rays that output an observation image corresponding to the viewpoint position of the three-dimensional object toward the viewpoint position, and a control unit that does not configure an observation image of any three-dimensional object in a region where the fields of view from the viewpoint positions of the respective users overlap when causing the display unit to display different three-dimensional objects for each viewpoint position of each user obtained from the captured images of the plurality of users. It has.

Effects of the Invention

[0007] According to the terminal device and the like in the present disclosure, it is possible to improve the quality of a stereoscopic image.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2A

Figure 2B

Figure 3A

Figure 3B

Embodiments for Carrying Out the Invention

[0009] FIG. 1 shows a configuration example of a virtual event providing system including a terminal device according to an embodiment. In the virtual event providing system 1, a plurality of terminal devices 12 and a server device 10 are connected to be able to communicate information with each other via a network 11. The virtual event providing system 1 is a system for providing an event in a virtual space that a user can participate in using the terminal device 12, that is, a virtual event. The virtual event is an event in which a plurality of users transmit information by speaking or the like in the virtual space, and each user is represented by a 3D model representing each of them, that is, a stereoscopic image.

[0010] The server device 10 belongs to, for example, a cloud computing system or other computing systems, and is a server computer that functions as a server for implementing various functions. The server device 10 has one or more communication interfaces, a storage device, a processor, or a dedicated circuit, and functions as a server computer. The server device 10 may be configured by two or more server computers that are connected to be able to communicate information and cooperate in operation. The server device 10 executes the transmission and reception of information necessary for providing a virtual event and information processing.

[0011] The terminal device 12 is an information processing device having a communication function, and is used by a user who participates in a virtual event provided by the server device 10. The terminal device 12 is, for example, an information processing terminal such as a smartphone or a tablet terminal, or an information processing device such as a personal computer.

[0012] The network 11 is, for example, the Internet, but includes an ad hoc network, a LAN (Local Area Network), a MAN (Metropolitan Area Network), or other networks or any combination thereof.

[0013] In the present embodiment, the terminal device 12 includes an imaging unit 117 that images a plurality of users, and a display / output unit 116 corresponding to a display unit that displays a three-dimensional object by configuring it with light rays that output an observation image corresponding to the viewpoint position of the three-dimensional object toward the viewpoint position, and a control unit 113. When the control unit 113 causes the display / output unit 116 to display different three-dimensional objects for each viewpoint position of each user obtained from the captured images of the plurality of users, in a region where the fields of view from the viewpoint positions of each user overlap, no observation image of any three-dimensional object is configured. When displaying three-dimensional images of different three-dimensional objects for each user, if the distance between users is short, an observation image of a three-dimensional object directed at another user may be mixed into the field of view of each user, and the clarity of the three-dimensional image that the user should view may be reduced. The terminal device 12 of the present embodiment can avoid the mixing of observation images directed at other users by not configuring an observation image of a three-dimensional object in the overlapping region of the fields of view of a plurality of users, and can prevent a decrease in the clarity of the three-dimensional image that should originally be viewed. That is, it is possible to improve the quality of the three-dimensional image.

[0014] The configuration of the terminal device 12 will be described in detail. In addition to the control unit 113, the imaging unit 117, and the display / output unit 116, the terminal device 12 includes a communication unit 111, a storage unit 112, and an input unit 115. Each unit is configured as follows, for example.

[0015] The communication unit 111 includes a communication module compatible with a wired or wireless LAN standard, a module compatible with a mobile communication standard such as LTE, 4G, 5G, etc. The terminal device 12 is connected to the network 11 via the communication unit 111 through a nearby router device or a base station of mobile communication, and performs information communication with the server device 10 etc. via the network 11.

[0016] The storage unit 112 includes, for example, a main storage device, an auxiliary storage device, or one or more semiconductor memories that function as cache memories, one or more magnetic memories, one or more optical memories, or a combination of at least two of these. The semiconductor memory is, for example, a RAM (Random Access Memory) or a ROM (Read Only Memory). The RAM is, for example, an SRAM (Static RAM) or a DRAM (Dynamic RAM). The ROM is, for example, an EEPROM (Electrically Erasable Programmable ROM). The storage unit 112 stores information used for the operation of the control unit 113 and information obtained by the operation of the control unit 113.

[0017] The control unit 113 includes, for example, one or more processors, one or more dedicated circuits, or a combination of these. The processor has, for example, one or more general-purpose processors such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), or a general-purpose processor specialized for a specific process, a dedicated processor such as a GPU (Graphics Processing Unit) specialized for a specific process. Alternatively, the control unit 113 may have one or more dedicated circuits such as an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The control unit 113 operates according to a control / processing program or according to an operation procedure implemented as a circuit, thereby comprehensively controlling the operation of the terminal device 12. Then, the control unit 113 transmits and receives various information to and from the server device 10 etc. via the communication unit 111, and executes the operation according to this embodiment.

[0018] The input unit 115 includes one or more input interfaces. The input interfaces include, for example, physical keys, capacitive keys, pointing devices, and touchscreens provided integrally with a display. The input interfaces also include a microphone for receiving voice input. Further, the input interfaces may include a scanner or camera for scanning an image code, and an IC card reader. The input unit 115 receives an operation for inputting information used for the operation of the control unit 113, and sends the input information to the control unit 113.

[0019] The display / output unit 116 includes one or more output interfaces for outputting information generated by the operation of the control unit 113. The output interfaces include, for example, a display and a speaker. The display is a light field display corresponding to the "display unit". The light field display irradiates light rays constituting an observation image corresponding to the viewpoint position of the three-dimensional object to be displayed to a plurality of viewpoint positions, enabling the user to visually recognize an observation image corresponding to their own viewpoint position and perceive a three-dimensional image of the three-dimensional object.

[0020] The imaging unit 117 includes a camera for imaging a captured image of a subject with visible light, and a distance measurement sensor for measuring the distance to the subject and obtaining a distance image. The camera generates a moving image composed of continuous captured images by imaging the subject, for example, at 15 to 30 frames per second. The distance measurement sensor includes a ToF (Time Of Flight) camera, LiDAR (Light Detection And Ranging), and a stereo camera, and generates a distance image of the subject including distance information. The imaging unit 117 sends the captured image and the distance image to the control unit 113.

[0021] The functions of the control unit 113 are realized by a processor included in the control unit 113 executing a control program. The control program is a program for causing the processor to function as the control unit 113. Also, some or all of the functions of the control unit 113 may be realized by a dedicated circuit included in the control unit 113. Further, the control program may be stored in a non-transitory recording and storage medium readable by the terminal device 12, and the terminal device 12 may read it from the medium.

[0022] In the present embodiment, the control unit 113 acquires a captured image and a distance image of the user of the terminal device 12 by the imaging unit 117, and collects the spoken voice of the user with the microphone of the input unit 115. The control unit 113 encodes the captured image and distance image of the user for generating a 3D model representing the user and the voice information for reproducing the voice of the user to generate encoded information. When encoding, the control unit 113 may perform arbitrary processing (such as resolution change and trimming) on the captured image and the like. The control unit 113 sends the encoded information to another terminal device 12 via the server device 10 by the communication unit 111.

[0023] Also, the control unit 113 receives the encoded information sent from another terminal device 12 via the server device 10 by the communication unit 111. When the control unit 113 decodes the encoded information received from another terminal device 12, it generates a 3D model representing the user using another terminal device 12 using the decoded information, and arranges the 3D model in the virtual space. The control unit 113 generates a virtual space image including the 3D model by rendering. The virtual space image includes observation images for each of a plurality of viewpoint positions for the light field display. That is, the control unit 113 renders a plurality of observation images. The control unit 113 outputs the virtual space image to the display / output unit 116. The display / output unit 116 displays the virtual space image by outputting the light rays constituting the plurality of observation images to each viewpoint position, and outputs the spoken voice based on the voice information of each user.

[0024] By the operations of the control unit 113 and the like, the user of the terminal device 12 can participate in the virtual event in real time and have conversations with other users.

[0025] FIGS. 2A and 2B are flowchart diagrams for explaining the operation procedures of the terminal device 12 related to the implementation of a virtual event. The server device 10 enables communication between the terminal devices 12 by mediating the communication of a plurality of terminal devices 12 on the network 11, and thus the virtual event is implemented.

[0026] The procedure of FIG. 2A relates to the operation procedure of the control unit 113 when each terminal device 12 sends information for generating 3D models of a plurality of users who use the terminal device 12.

[0027] In step S201, the control unit 113 causes the imaging unit 117 to capture a visible light image of the user and acquire a distance image at an arbitrarily set frame rate, and causes the input unit 115 to collect the voice of the self-user's speech. The control unit 113 acquires the captured image by visible light and the distance image from the imaging unit 117, and acquires voice information from the input unit 115.

[0028] In step S202, the control unit 113 derives the respective viewpoint positions for each user. The viewpoint position of the user is specified by the spatial coordinates of an arbitrary part of the user's face capable of image recognition. The arbitrary part is, for example, either eye or the midpoint between both eyes, etc. The control unit 113 recognizes the user and the part of the face by image processing such as pattern matching on the captured image. Also, the control unit 113 derives the spatial coordinates of, for example, the user's eye with respect to the position of the camera of the imaging unit 117 based on the distance image. The control unit 113 derives the viewpoint position of the user based on the obtained spatial coordinates. Also, the control unit 113 derives the visual field of each user based on the viewpoint position and the line of sight. The line of sight is obtained by recognizing the convergence angle of the eyeballs through image recognition. The control unit 113 derives, for example, an arbitrary spatial region centered on the intersection point of the line of sight as the visual field. The spatial region as the visual field is derived using, for example, the distance from the viewpoint position to the intersection point of the line of sight, the general effective visual field angle of a human, and the general depth of field of the eye. Information for obtaining the visual field is stored in the storage unit 112 in advance, and the control unit 113 derives the visual field using such information. In one example, for instance, when the intersection point of the user's line of sight is located about 60 to 80 cm in front of the viewpoint position, a space with a width of 20 to 30 cm, a height of 10 to 20 cm, and a depth of about 10 cm centered on the intersection point of the line of sight corresponds to the visual field.

[0029] In step S203, the control unit 113 encodes the captured image, the distance image, and the audio information to generate encoded information.

[0030] In step S204, the control unit 113 packets the encoded information by the communication unit 111 and sends it to the server device 10 toward the other terminal device 12.

[0031] When the control unit 113 acquires information input in response to an operation for interrupting imaging and sound collection or an operation for exiting a virtual event (Yes in S205), it ends this processing procedure. While the control unit 113 does not acquire information corresponding to an operation for interruption or exit (No in S205), it executes steps S201 to S204 to derive the viewpoint positions of each user and to transmit information for generating a 3D model representing each user and information for outputting sound.

[0032] The procedure of FIG. 2B relates to the operation procedure of the control unit 113 when the terminal device 12 outputs a virtual event image and the voice of another user using another terminal device 12. When the control unit 113 receives, via the server device 10, a packet transmitted by another terminal device 12 executing the procedure of FIG. 2A, it executes steps S211 to S213.

[0033] In step S211, the control unit 113 decodes the encoded information included in the packet received from another terminal device 12 to acquire a captured image, a distance image, and voice information.

[0034] In step S212, the control unit 113 generates a 3D model representing another user based on the captured image and the distance image. When generating the 3D model, the control unit 113 generates a polygon model using the distance image of another user and performs texture mapping using the captured image of another user on the polygon model, thereby generating a 3D model of another user. However, the generation of the 3D model is not limited to the example shown here, and any method can be adopted.

[0035] When receiving information from the terminal devices 12 of a plurality of other users, the control unit 113 executes steps S211 to S212 for each of the other terminal devices 12 to generate a 3D model for each other user.

[0036] In step S213, the control unit 113 arranges 3D models representing other users in the virtual space where the virtual event is held. The control unit 113 arranges the generated 3D models of other users at the coordinates in the virtual space. The storage unit 112 stores the coordinate information of the virtual space and the information of the coordinates where each 3D model should be arranged. For example, the 3D models of other users are assigned arrangements according to the order in which other users log in to the virtual event. Alternatively, it is also possible for the 3D model to move within the virtual space by the operation of other users. In that case, the position corresponding to the operation of other users is assigned to the 3D model.

[0037] In step S214, the control unit 113 determines the three-dimensional objects to be displayed for each user. For example, the control unit 113 selectively displays, for each user, the 3D model of the other user arranged at the position closest to the viewpoint position of each user. The display target may be a 3D model of a real or fictional installation object other than other users. The storage unit 112 stores information associating the spatial coordinates in the virtual space with the spatial coordinates in the real space when the virtual space image is displayed by the light field display. Using this information, the control unit 113 determines, as the object to be displayed, the three-dimensional object in the virtual space closest to the viewpoint position of each user in the real space.

[0038] In step S215, the control unit 113 renders and generates a virtual space image obtained by imaging the 3D models arranged in the virtual space from the viewpoint positions of each user. The control unit 113 renders and generates, for each user, a virtual space image including the three-dimensional objects to be displayed determined in step S214. The virtual space image includes a plurality of observation images corresponding to the number of assumed viewpoint positions. The number of viewpoint positions corresponds to the number of angles at which the light field display can emit light rays and is stored in advance in the storage unit 112. The control unit 113 uses this information to generate observation images corresponding to the number of viewpoint positions.

[0039] In step S216, the control unit 113 causes the display / output unit 116 to display the virtual space image and output sound. That is, the control unit 113 outputs information for displaying the virtual space image to the display / output unit 116. The display / output unit 116 displays the virtual space image by outputting light rays constituting the observation image toward each viewpoint position, and outputs sound.

[0040] Here, the display of the virtual space image will be described with reference to FIGS. 3A and 3B. FIG. 3A is a plan view schematically showing the display mode of the virtual space image. Here, a configuration in a plan view is shown in which a light field display 30 outputs light rays 35 for constituting an element image of the virtual space image from pixel units 32 (when referring to each individually, pixel units 32a, 32b, 32c, and 32d) arranged on its display surface 31 toward users 300 and 301. FIG. 3B shows the configuration of the pixel unit 32.

[0041] As shown in FIG. 3B, each pixel unit 32 has a microlens 34 and a plurality of corresponding pixels 33. Each pixel 33 includes a light emitting element that outputs a light ray 35 corresponding to a pixel value instructed by the control unit 113. The light ray 35 is irradiated by the microlens 34 in a plurality of angles, that is, toward a plurality of assumed viewpoint positions. Here, schematically, the light ray 35 shown by a solid line constitutes an observation image of a 3D model 36 of another user to be displayed toward the user 300, and the light ray 35 shown by a dotted line constitutes an observation image of a 3D model 37 of another user to be displayed toward the user 301. The observation image corresponds to an image visually recognized by the user at the viewpoint position corresponding to each light ray 35.

[0042] As shown in FIG. 3A, the control unit 113 outputs light rays 35 for constructing observation images of the 3D models 36 and 37, respectively, to the individual pixel units 32a to 32d of the light field display 30 so as to display the 3D models 36 and 37 in the respective visual fields 302 and 303 of the users 300 and 301. At this time, the control unit 113 selectively outputs, to the pixel 33, the light ray 35a for constructing the observation image of the 3D model 36 directed to the visual field 302 of the user 300 and the light ray 35b for constructing the observation image of the 3D model 37 directed to the visual field 303 of the user 301 according to the positions of the pixel units 32a to 32d. However, in the present embodiment, the control unit 113 is configured not to construct an observation image of any 3D model in the overlapping region 304 of the respective visual fields 302 and 303 of the users 300 and 301. For example, the control unit 113 identifies the pixel 33 that outputs the light ray 305 to the overlapping region 304, and stops the light emission or outputs a light ray for a low-luminance masking image. By the operations of the control unit 113 and the display / output unit 116, it is possible to avoid the inconvenience that observation images other than the stereoscopic images that should be originally displayed are mixed in the respective visual fields 302 and 303 of the users 300 and 301, thereby preventing the visual recognition of the original stereoscopic images. Note that the number of pixel units 32, the number of light rays 35, etc. shown here are simplified examples, and it goes without saying that none of them are limited to the examples shown here.

[0043] By repeatedly executing steps S211 to S216 by the control unit 113, each user can listen to the voice of the speech of another user while watching a moving image of a virtual space image including the 3D model of the nearest other user.

[0044] By executing the procedures of FIGS. 2A and 2B in parallel, for example, in a time-division manner, the user can use the terminal device 12 to interact with other users while viewing each other's 3D models in a virtual event. At this time, the control unit 113 can improve the image of the stereoscopic image by not constructing an observation image in the region where the respective visual fields overlap according to the viewpoint positions of a plurality of users.

[0045] In the above, the embodiments have been described based on the drawings and examples. However, it should be noted that those skilled in the art can easily make various modifications and alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are included within the scope of the present disclosure. For example, the functions included in each means, each step, etc. can be rearranged so as not to be logically inconsistent, and a plurality of means, steps, etc. can be combined into one or divided.

Description of Reference Numerals

[0046] 1 Virtual event providing system 10 Server device 11 Network 12 Terminal device 101, 111 Communication unit 102, 112 Storage unit 103, 113 Control unit 115 Input unit 116 Display / output unit 117 Imaging unit

Claims

A system having a pair of terminal devices capable of communicating with each other, each terminal device comprising: an imaging unit that images a plurality of users; a display unit that displays the three-dimensional object so that each user can visually recognize it by configuring it with light rays that output an observation image of the three-dimensional object at a predetermined position toward the viewpoint position corresponding to each user's viewpoint position with respect to the predetermined position; a control unit that, when causing the display unit to display different observation images for each viewpoint position of each user obtained from the captured images of the plurality of users, does not configure any observation images in a region where the fields of view from the viewpoint positions of the respective users overlap; and having one of the pair of terminal devices sends information for generating a 3D model representing at least one of the plurality of users corresponding to the one terminal device to the other terminal device of the pair of terminal devices, and causes the other terminal device to display the 3D model as the three-dimensional object. System. **Claim 2** In claim 1, the control unit does not output light rays to the viewpoint positions in the region where the fields of view from the viewpoint positions of the respective users overlap to the display unit. System. **Claim 3** In claim 1, the control unit outputs light rays for a low-luminance masking image to the viewpoint positions in the region where the fields of view from the viewpoint positions of the respective users overlap to the display unit. System.

Citation Information

Patent Citations

  • Directivity-expressing display device

    JP2008262107A

  • Rendering method and apparatus for plurality of users

    JP2017038367A

  • System and method for implementing a viewer-specific image perception adjustment within a defined view zone, and vision correction system and method using same

    US20220394234A1

  • Information processing device, information processing system, method for controlling information processing device, and method for setting parameter

    WO2017094543A1

  • Light field displays incorporating eye trackers and methods for generating views for a light field display using eye tracking information

    WO2021087450A1