Web-based video conferencing virtual environment with controllable avatars and its applications

The method of texture-mapping video streams onto avatars in a 3D virtual space addresses the limitations of traditional video conferencing by enhancing social interaction and efficiently managing resources for large groups, providing a more immersive and interactive virtual meeting experience.

JP7717123B2Active Publication Date: 2025-08-01KATMAI TECH INC
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
JP2023117467
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-11
Filing Date
2023-07-19
Publication Date
2025-08-01
Estimated Expiration
2041-10-20

AI Technical Summary

Technical Problem

Traditional video conferencing lacks the sense of place and social interaction experienced in physical meetings, leading to a loss of experiential and social connections, and is limited by network bandwidth and computing hardware, making it difficult to handle large numbers of participants efficiently.

Method used

A method for video conferencing that uses a three-dimensional virtual space where participants' video streams are texture-mapped onto avatars, with audio streams adjusted for spatial positioning and volume, and bandwidth allocation based on distance, enabling private conversations and efficient resource management.

Benefits of technology

Enhances social interaction and spatial awareness in virtual meetings, allowing for private conversations and efficient handling of large participant counts by optimizing bandwidth and computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007717123000001
    Figure 0007717123000001
  • Figure 0007717123000002
    Figure 0007717123000002
  • Figure 0007717123000003
    Figure 0007717123000003
Patent Text Reader

Abstract

To allow video avatars (102A, 102B) to navigate within a virtual environment.SOLUTION: A system has a presented mode that allows for a presentation stream to be texture mapped to a presenter screen (104A, 104B) situated within a virtual environment. The relative left-right sound is adjusted to provide sense of an avatar's position in a virtual space. The sound is further adjusted based on the area where the avatar is located and where the virtual camera is located. Video stream quality is adjusted based on relative position in a virtual space. Three-dimensional modeling is available inside a virtual video conferencing environment.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is a continuation of U.S. Utility Patent Application No. 17 / 075,338, filed October 20, 2020, now issued as U.S. Patent No. 10,979,672, issued April 13, 2021; U.S. Utility Patent Application No. 17 / 198,323, filed March 11, 2021; U.S. Utility Patent Application No. 17 / 075,362, filed October 20, 2020, now issued as U.S. Patent No. 11,095,857, issued August 17, 2021; and U.S. Utility Patent Application No. 17 / 075,362, filed October 20, 2020, now issued as U.S. Patent No. 10,952,006, issued March 16, 2021. This application claims priority to U.S. Utility Patent Application No. 17 / 075,390, filed October 20, 2020, U.S. Utility Patent Application No. 17 / 075,408, filed October 20, 2020, now issued as U.S. Patent No. 11,070,768, issued July 20, 2021, U.S. Utility Patent Application No. 17 / 075,428, filed October 20, 2020, now issued as U.S. Patent No. 11,076,128, issued July 27, 2021, and U.S. Utility Patent Application No. 17 / 075,454, filed October 20, 2020, the contents of each of which are incorporated herein by reference in their entirety.

[0002] background Field

[0002] The field generally relates to videoconferencing. [Background technology]

[0003] Related technologies

[0003] Video conferencing involves the transmission and reception of audiovisual signals by multiple users in different locations for real-time communication among people. Video conferencing is widely available on many computing devices from a variety of services, including the ZOOM service available from Zoom Communications Inc. located in San Jose, CA. Some video conferencing software, such as the FaceTime application available from Apple Inc. located in Cupertino, CA, is standard on mobile devices.

[0004]

[0004] Generally, these applications operate by displaying the video of other conference participants and outputting audio. When there are multiple participants, the screen can be divided into several rectangular frames, each displaying the video of one participant. These services may also operate by having a larger frame that presents the video of the person speaking. When different individuals speak, the frame switches between speakers. The application captures video from a camera integrated into the user's device and audio from a microphone integrated into the user's device. The application then transmits that audio and video to other applications running on other users' devices.

[0005]

[0005] Many of these video conferencing applications have a screen sharing feature. When a user specifies the sharing of their screen (or a portion of their screen), a stream is sent to the devices of other users along with the content of the screen. In some cases, other users can even control what is on that user's screen. In this way, users can collaborate on projects or give presentations to other meeting participants.

[0006]

[0006] In recent years, video conferencing technology has become increasingly important. Due to fears of illness, particularly the spread of COVID-19, many workplaces, trade shows, meetings, conferences, schools, and places of worship have been closed or discouraged from participation. Virtual meetings using video conferencing technology are increasingly replacing physical meetings. Additionally, this technology offers advantages over physically meeting to avoid business trips and commuting.

[0007]

[0007] However, in many cases, using this video conferencing technology results in a loss of the sense of place. There is an experiential aspect to physically meeting face-to-face in the same location, which is lost when the meeting is conducted virtually. There is also a social aspect where users can pose and see their peers. This sense of experience is important in creating relationships and social connections. In terms of traditional video conferencing, this sense is currently lacking.

[0008]

[0008] Furthermore, when a meeting starts with a few participants, further problems arise with these video conferencing technologies. In a physical meeting, people can engage in small talk. They can speak in voices audible only to those nearby. In some cases, they can even have private conversations in the context of a larger meeting. However, when using a virtual meeting and multiple people speak simultaneously, the software mixes the two audio streams approximately equally, causing participants to interrupt each other. Therefore, when multiple people are involved in a virtual meeting, private conversations are impossible, and the interaction tends to be in the form of one person speaking to many people instead. Here too, virtual meetings miss the opportunity for participants to create social connections and communicate more efficiently to expand their networks.

[0009]

[0009] Furthermore, due to network bandwidth and computing hardware limitations, when many streams occur in a meeting, the performance of many video conferences begins to slow down. Many computing devices are equipped to handle video streams from 2 to 3 participants and are not equipped to handle video streams from more than 12 participants. In many schools that are virtually run overall, a class of 25 can significantly slow down school-issued computing devices.

[0010]

[0010] Massively multiplayer online games (MMOGs or MMOs) can generally handle significantly more than 25 participants. In these games, often hundreds or thousands of players are on a single server. In MMOs, players can often navigate their avatars throughout the virtual world. In these MMOs, users may be able to talk to each other or send messages to each other. Examples include the ROBLOX game available from Roblox Corporation of San Mateo, CA and the MINECRAFT game available from Mojang Studios of Stockholm, Sweden.

[0011]

[0011] There are also limitations regarding social interaction in getting naked avatars to interact with each other. These avatars usually cannot convey the expressions that people often make unintentionally. These expressions are observable in video conferences. Some published reports may describe videos placed on avatars in virtual worlds. However, these systems typically require dedicated software and have other limitations that restrict their usefulness.

Summary of the Invention

Problems to be Solved by the Invention

[0012]

[0012] An improved method for video conferencing is needed.

Means for Solving the Problems

[0013] Summary

[0013] In one embodiment, the device enables a video conference between a first user and a second user. The device includes a processor coupled to a memory, a display screen, a network interface, and a web browser. The network interface is configured to receive (i) data specifying a three-dimensional virtual space, (ii) a position and a direction in the three-dimensional virtual space, the position and the direction being input by the first user, and (iii) a video stream captured from a camera of the first user's device. The camera of the first user is positioned to capture a photographic image of the first user. The web browser is implemented by the processor and is configured to download a web application from a server and execute the web application. The web application includes a texture mapper and a renderer. The texture mapper is configured to texture map the video stream onto a three-dimensional model of an avatar. The renderer is configured to render the three-dimensional virtual space including the texture-mapped three-dimensional model of the avatar disposed at the position and oriented in the direction for display to the second user from a viewpoint of a virtual camera of the second user. By managing texture mapping within the web application, the embodiment avoids the need to install dedicated software.

[0014]

[0014] In one embodiment, the computer-implemented method enables a presentation in a virtual conference including a plurality of participants. In this method, data specifying a three-dimensional virtual space is received. The position and orientation in the three-dimensional virtual space are also received. The position and orientation are input by a first participant among the plurality of participants in the conference. Finally, a video stream captured from the camera of the first participant's device is received. The camera is positioned to capture a photographic image of the first participant. The video stream is texture mapped onto a three-dimensional model of an avatar. In addition, a presentation stream from the first participant's device is received. The presentation stream is texture mapped onto a three-dimensional model of a presentation image. Finally, a three-dimensional virtual space having the texture-mapped avatar and the texture-mapped presentation screen is rendered for display to a second participant among the plurality of participants from the perspective of the second participant's virtual camera. In this way, the embodiment enables a presentation in a social conference environment.

[0015]

[0015] In one embodiment, the computer-implemented method provides audio to a virtual conference including a plurality of participants. In this method, a three-dimensional virtual space including an avatar having the texture-mapped video of a second user is rendered for display to a first user from the perspective of the first user's virtual camera. The virtual camera is at a first position in the three-dimensional virtual space, and the avatar is at a second position in the three-dimensional virtual space. An audio stream from the microphone of the second user's device is received. The microphone is positioned to capture the speech of the second user. The volume of the received audio stream is adjusted to provide a sense that the second position is relative to the first position in the three-dimensional virtual space by identifying a left audio stream and a right audio stream. The left audio stream and the right audio stream are output in stereo and played back to the first user.

[0016]

[0016] In one embodiment, a computer-implemented method provides audio for a virtual conference. In this method, a three-dimensional virtual space including an avatar having the second user's texture-mapped video is rendered for display to a first user from the perspective of the first user's virtual camera. The virtual camera is at a first position in the three-dimensional virtual space, and the avatar is at a second position in the three-dimensional virtual space. An audio stream from a microphone of the second user's device is received. It is determined whether the virtual camera and the avatar are arranged in the same area among a plurality of areas. If it is determined that the virtual camera and the avatar are not arranged in the same area, the audio stream is attenuated. The attenuated audio stream is output and played back to the first user. In this way, the embodiment enables private conversations and small talk in a virtual video conferencing environment.

[0017]

[0017] In one embodiment, a computer-implemented method efficiently streams video for a virtual conference. In this method, the distance between a first user and a second user in a virtual conference space is determined. A video stream captured from a camera of the first user's device is received. The camera is positioned to capture a photographic image of the first user. The resolution or bitrate of the video stream is reduced based on the determined distance such that the closer the distance, the higher the resolution compared to when the distance is far. The video stream is transmitted at the reduced resolution or bitrate to the second user's data for display to the second user within the virtual conference space. The video stream should be texture-mapped to the first user's avatar for display to the second user within the virtual conference space. In this way, the embodiment efficiently allocates bandwidth and computing resources even when there are a large number of conference participants.

[0018]

[0018] In one embodiment, the computer-implemented method enables modeling in a virtual video conference. In this method, a mesh representing a three-dimensional model of an object, which is a three-dimensional model of a virtual environment, and a video stream from a participant in the virtual video conference are received. The video stream is texture mapped to an avatar that can be manipulated by the participant. The texture-mapped avatar and the mesh representing the three-dimensional model of the object in the virtual environment are rendered for display.

[0019]

[0019] Embodiments of systems, devices, and computer program products are also disclosed.

[0020]

[0020] Further embodiments, features, and advantages of the present invention, as well as the structure and operation of various embodiments, will be described in detail below with reference to the accompanying drawings.

[0021] Brief Description of the Drawings

[0021] The accompanying drawings are incorporated herein and form a part of the present invention, show the present disclosure, and together with the description further function to explain the principles of the present disclosure so that those skilled in the art can make and use the present disclosure.

Brief Description of the Drawings

[0022]

Figure 1

[0022] FIG. shows an example of an interface that provides a video conference in a virtual environment in which a video stream is mapped to an avatar.

Figure 2

[0023] FIG. shows a three-dimensional model used to render a virtual environment having an avatar for a video conference.

Figure 3

[0024] FIG. shows a system that provides a video conference in a virtual environment.

Figure 4A

[0025] FIG. shows how data is transferred between various components of the system of FIG. 3 to provide a video conference.

Figure 4B

[0025] To provide a video conference, it shows how data is transferred between various components of the system in FIG. 3.

Figure 4C

[0025] To provide a video conference, it shows how data is transferred between various components of the system in FIG. 3.

Figure 5

[0026] It is a flowchart showing a method of adjusting the left and right volumes relatively to provide a sense of position in a virtual environment during a video conference.

Figure 6

[0027] It is a chart showing how the volume decreases as the distance between avatars increases.

Figure 7

[0028] It is a flowchart showing a method of adjusting the relative volume to provide different volume areas in a virtual environment during a video conference.

Figure 8A

[0029] It is a diagram showing different volume areas in a virtual environment during a video conference.

Figure 8B

[0029] It is a diagram showing different volume areas in a virtual environment during a video conference.

Figure 9A

[0030] It is a diagram showing the traversal of volume areas in a virtual environment during a video conference.

Figure 9B

[0030] It is a diagram showing the traversal of volume areas in a virtual environment during a video conference.

Figure 9C

[0030] It is a diagram showing the traversal of volume areas in a virtual environment during a video conference.

Figure 10

[0031] It shows the interaction with a three-dimensional model in a three-dimensional virtual environment.

Figure 11

[0032] It shows the presentation screen sharing in a three-dimensional virtual environment used for a video conference.

Figure 12

[0033] It is a flowchart showing a method of allocating available bandwidth based on the relative positions of avatars within a three-dimensional virtual environment.

Figure 13

[0034] A chart showing how the priority value can decrease as the distance between avatars increases.

Figure 14

[0035] A chart showing how the band wave allocated based on the relative priority can be changed.

Figure 15

[0036] A diagram showing the components of a device used to provide a video conference in a virtual environment.

Best Mode for Carrying Out the Invention

[0023]

[0037] The drawing in which an element first appears is typically indicated by one or more digits at the left end in the corresponding reference number. In the drawings, like reference numbers may indicate the same element or elements that are functionally the same.

[0024] Detailed Description Video Conferencing Using Avatars in a Virtual Environment

[0038] FIG. 1 is a diagram showing an example interface 100 for providing a video conference in a virtual environment in which a video stream is mapped to an avatar.

[0025]

[0039] Interface 100 can be displayed to participants in a video conference. For example, Interface 100 can be rendered for display to participants and can be updated in real time as the video conference progresses. The user can control the orientation of the user's virtual camera, for example, using keyboard input. In this way, the user can navigate around the virtual environment. In one embodiment, different inputs can change the X and Y positions, pan angle, and tilt angle of the virtual camera in the virtual environment. In a further embodiment, the user can use an input to change the height (Z coordinate) or yaw of the virtual camera. In a further embodiment, the user can input an input to simulate gravity by causing the virtual camera to "hop" up while the virtual camera returns to its original position. Inputs available for navigating the virtual camera can include, for example, the WASD keyboard keys to move the virtual camera forward, backward, left, or right on the X-Y plane, the space bar key to "hop" the virtual camera, and keyboard and mouse inputs such as mouse movement to specify changes in the pan angle and tilt angle.

[0026]

[0040] Interface 100 includes avatars 102A and 102B, each avatar representing a different participant in the video conference. Avatars 102A and 102B have video streams 104A and 104B from the devices of the first and second participants texture mapped thereon. A texture map is an image that is applied (mapped) to the surface of a shape or polygon. Here, the image is each frame of the video. The camera devices that capture video streams 104A and 104B are positioned to capture the face of each participant. In this way, the avatars are texture mapped and move the face image as the participants in the meeting listen to the conversation.

[0027]

[0041] Similar to the way the virtual camera is controlled by the user browsing interface 100, the locations and orientations of avatars 102A and 102B are controlled by each participant represented by the avatar. Avatars 102A and 102B are three-dimensional models represented by a mesh. Each of avatars 102A and 102B may have the name of the participant under the avatar.

[0028]

[0042] Each of avatars 102A and 102B is controlled by various users. Each avatar can be positioned at a point corresponding to the location where the avatar's own virtual camera is placed within the virtual environment. In exactly the same way that the user browsing interface 100 can move around the virtual camera, various users can move around each of avatars 102A and 102B.

[0029]

[0043] The virtual environment rendered on interface 100 includes a background image 120 and a three-dimensional model 118 of an arena. The arena can be the venue or building where the video conference is to be held. The arena can include a floor area partitioned by walls. The three-dimensional model 118 can include a mesh and textures. Other ways of mathematically representing the surface of the three-dimensional model 118 may similarly be possible. For example, polygon modeling, curve modeling, and digital sculpting may be possible. For example, the three-dimensional model 118 can be represented by voxels, splines, geometric primitives, polygons, or any other possible representation in three-dimensional space. The three-dimensional model 118 can also include specifications of light sources. The light sources can include, for example, point light sources, directional light sources, spotlight light sources, and ambient light sources. The objects can also have specific properties that describe how light is reflected. In the example, the properties can include diffuse illumination interaction, ambient illumination interaction, and spectral illumination interaction.

[0030]

[0044] In addition to the arena, the virtual environment can include various other three-dimensional models that represent different components of the environment. For example, the three-dimensional environment can include a decoration model 114, a speaker model 116, and a presentation screen model 122. Similar to model 118, these models can be represented using any mathematical method that represents geometric surfaces in three-dimensional space. These models may be separate from model 118 or may be combined into a single representation of the virtual environment.

[0031]

[0045] Decoration models such as model 114 function to enhance realism and increase the aesthetic appeal of the arena. Speaker model 116 can virtually emit sounds such as presentations and background music, as described in more detail with respect to FIGS. 5 and 7. Presentation screen model 122 can function to provide an outlet for exemplifying a presentation. A video of a presenter or presentation screen sharing can be texture mapped onto presentation screen model 122.

[0032]

[0046] Button 108 can provide the user with a list of participants. In one example, after the user selects button 108, the user can chat with other participants by sending text messages individually or as a group.

[0033]

[0047] Button 110 can enable the user to change the attributes of the virtual camera used for rendering the interface 100. For example, the virtual camera can have a field of view that specifies the angle at which data is rendered for display. The modeling of the data within the camera's field of view is rendered, while the modeling of the data outside the camera's field of view may not be rendered. By default, the field of view of the virtual camera can be set anywhere between 60° and 110°, which is compatible with a wide-angle lens and human vision. However, when button 110 is selected, the virtual camera can increase its field of view beyond 170°, which is compatible with a fish-eye lens. This enables the user to potentially have a wider peripheral awareness within the virtual environment.

[0034]

[0048] Finally, button 112 enables the user to exit the virtual environment. When button 112 is selected, the device can be signaled to stop displaying the avatar corresponding to the user who was previously viewing interface 100 to the devices belonging to other participants.

[0035]

[0049] In this way, a video conference is conducted using the interface virtual 3D space. Any user can control the avatar, and the user can control the avatar to move around, look around, jump, or perform other actions such as changing position or orientation. The virtual camera shows the user the virtual 3D environment and other avatars. The avatars of other users have a virtual display that shows the user's webcam image as an integral part.

[0036]

[0050] By giving the user a sense of space and enabling users to see each other's faces, the embodiments provide a more social experience than traditional web conferencing or traditional MMO games. The more social experience has various applications. For example, it can be used in online shopping. For example, the interface 100 can be a virtual supermarket, a place of worship, a fair, B2B sales, B2C sales, schooling, a restaurant or cafeteria, a product release, a construction site visit (e.g., for architects, engineers, contractors), an office space (e.g., where people work virtually "at their desks"), remote control of machinery (ships, vehicles, airplanes, submarines, drones, drilling equipment, etc.), a factory / facility control room, a medical procedure, garden design, a guided virtual bus tour, a music event (e.g., a concert), a lecture (e.g., a TED talk), a meeting of a political group, a board meeting, an underwater survey, an investigation of a hard-to-reach location, training for emergencies (e.g., a fire), cooking, shopping (payment and delivery), virtual art and crafts (e.g., painting and pottery), a wedding, a funeral, a baptism, remote sports training, counseling, dealing with phobias (e.g., face-to-face therapy), a fashion show, an amusement park, home decoration, sports viewing, e-sports viewing, viewing a performance captured using a 3D camera, playing board games and role-playing games, walking on / through medical images, viewing geological data, language learning, meetings in a space for visually impaired people, meetings in a space for hearing impaired people, participation in events by people who normally cannot walk or stand, presentation of news or weather, talk shows, book signings, voting, MMO, purchase / sale of virtual locations (Linden Research, Inc. located in San Francisco, CAAvailable in some MMOs such as the SECOND LIFE game (etc.), flea markets, garage sales, travel agencies, banks, archives, computer process management, fencing / sword fighting / martial arts, reenactment videos (e.g., crime scene and / or accident reenactment), rehearsal of real events (e.g., weddings, presentations, shows, spacewalks), evaluation or viewing of real events captured by a 3D camera, livestock shows, zoos, life experiences as tall people / short people / visually impaired people / hearing impaired people / white people / black people (e.g., virtual world modified video streams or still images to simulate the perspective from which the user wants to experience the reaction), job interviews, game shows, interactive fiction (e.g., murder mysteries), virtual fishing, virtual sailing, psychological surveys, behavior analysis, virtual sports (e.g., climbing / bouldering), lighting control at home or other places (home automation), memory palaces, archaeology, gift shops, virtual visits to make it more comfortable when customers actually visit, virtual medical treatments to explain procedures and make people feel more comfortable, and virtual exchanges / financial markets / stock markets (e.g., integrating real-time data and video feeds for real-time trading and analysis in a virtual world), virtual places where people need to go as part of their jobs so that they can actually meet with each other organizationally (e.g., if you want to create an invoice, you can only do so from within the virtual place), and augmented reality that projects a person's face onto an AR headset (or helmet) so that facial expressions (e.g., useful for military, law enforcement, firefighters, special operations) can be seen, and has uses in providing reservations (e.g., for specific villas / cars / etc.).

[0037]

[0051] FIG. 2 is a diagram 200 showing a three-dimensional model used to render a virtual environment having avatars for a video conference. Similar to what is shown in FIG. 1, the virtual environment here includes a three-dimensional arena 118 and various three-dimensional models including three-dimensional models 114 and 122. Also as shown in FIG. 1, FIG. 200 includes avatars 102A and 102B that navigate various parts of the virtual environment.

[0038]

[0052] As described above, the interface 100 in FIG. 1 is rendered from the perspective of a virtual camera. The virtual camera is shown as virtual camera 204 in FIG. 200. As previously described, the user browsing interface 100 in FIG. 1 can control the virtual camera 204 and navigate the virtual camera in three-dimensional space. The interface 100 is consistently updated according to the new position of the virtual camera 204 and any changes to the models within the view of the virtual camera 204. As described above, the field of view of the virtual camera 204 can be a frustum that is at least partially defined by the horizontal and vertical fields of view of the viewing angle.

[0039]

[0053] As described above with respect to FIG. 1, the background image or texture can define at least a portion of the virtual environment. The background image can capture the sides of the virtual environment that are intended to appear distant. The background image can be texture mapped to the sphere 202. The virtual camera 204 can be at the origin of the sphere 202. In this way, distant features of the virtual environment can be rendered efficiently.

[0040]

[0054] In other embodiments, other shapes can be used instead of the sphere 202 and texture mapped to the background image. In various alternative embodiments, the shape can be cylindrical, cubic, rectangular prism, or any other three-dimensional geometry.

[0041]

[0055] FIG. 3 is a diagram showing a system 300 for providing a video conference in a virtual environment. The system 300 includes a server 302 coupled to devices 306A and 306B via one or more networks 304.

[0042]

[0056] Server 302 provides a service for connecting a video conference session between devices 306A and 306B. As described in more detail below, Server 302 communicates notifications to the devices of the conference participants (e.g., devices 306A and 306B) when a new participant joins the conference and when an existing participant leaves the conference. Server 302 communicates messages that describe the position and orientation in the three-dimensional virtual space of the virtual cameras of each participant in the three-dimensional virtual space. Server 302 also communicates video and audio streams between each participant's devices (e.g., devices 306A and 306B). Finally, Server 302 stores data that describes data specifying the three-dimensional virtual space and transmits it to each of devices 306A and 306B.

[0043]

[0057] In addition to the data required for the virtual conference, Server 302 may provide executable information that instructs devices 306A and 306B on how to render the data in order to provide an interactive conference.

[0044]

[0058] Server 302 responds in response to requests. Server 302 can be a web server. A web server is software and hardware that uses HTTP (HyperText Transfer Protocol) and other protocols to respond to client requests made via the World Wide Web. The main job of a web server is to display website content through the storage, processing, and delivery of web pages to users.

[0045]

[0059] In an alternative embodiment, the communication between devices 306A and 306B is not done through Server 302 but is done on a peer-to-peer basis. In that embodiment, data that describes the location and orientation of each participant, notifications regarding new and existing participants, and one or more of the video streams and audio streams of each participant are communicated directly between devices 306A and 306B rather than through Server 302.

[0046]

[0060] Network 304 enables communication between various devices 306A and 306B and server 302. Network 304 can be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless wide area network (WWAN), a metropolitan area network (MAN), a part of the Internet, a part of the public switched telephone network (PSTN), a cellular telephone network, a wireless network, a WiFi network, a WiMax network, any other type of network, or any combination of two or more such networks.

[0047]

[0061] Devices 306A and 306B are each device of each participant in the virtual meeting. Devices 306A and 306B each receive data necessary to conduct the virtual meeting and render data necessary to provide the virtual meeting. As described in more detail below, devices 306A and 306B include a display for presenting the rendered meeting information, an input that enables the user to control the virtual camera, a speaker (such as a headset) that provides audio to the user for the meeting, a microphone for capturing the user's voice input, and a camera positioned to capture video of the user's face.

[0048]

[0062] Devices 306A and 306B can be any type of computing device, including a laptop, a desktop, a smartphone or tablet computer, or a wearable computer (such as a smartwatch, or an augmented reality, or virtual reality headset, etc.).

[0049]

[0063] Web browsers 308A and 308B can search for network resources (such as web pages) addressed by a link identifier (e.g., a Uniform Resource Locator or URL) and present the network resources for display. In particular, web browsers 308A and 308B are software applications for accessing information on the World Wide Web. Typically, web browsers 308A and 308B make this request using the Hypertext Transfer Protocol (HTTP or HTTPS). When a user requests a web page from a particular website, the web browser searches for the required content from the web server, interprets and executes the content, and then displays the page on the displays of devices 306A and 306B shown as client / peer meeting applications 308A and 308B. In the example, the content may have client-side scripts such as HTML and JavaScript. Once displayed, the user can enter information and make selections on the page, whereby web browsers 308A and 308B make further requests.

[0050]

[0064] The conference applications 310A and 310B can be web applications that are downloaded from the server 302 and configured to be executed by each of the web browsers 308A and 308B. In one embodiment, the conference applications 310A and 310B can be JavaScript applications. In one example, the conference applications 310A and 310B can be written in a high-level language such as the TypeScript language and translated or compiled into JavaScript. The conference applications 310A and 310B are configured to interact with the WebGL JavaScript application programming interface. They can have control code specified in JavaScript and shader code written in the OpenGL ES shading language (GLSL ES). Using the WebGL API, the conference applications 310A and 310B may be able to utilize the graphics processing units (not shown) of the devices 306A and 306B. Furthermore, OpenGL rendering of interactive two-dimensional and three-dimensional graphics without using a plugin.

[0051]

[0065] The conference applications 310A and 310B receive from the server 302 data describing the positions and orientations of other avatars and three-dimensional modeling information describing the virtual environment. In addition, the conference applications 310A and 310B receive from the server 302 video streams and audio streams of other conference participants.

[0052]

[0066] The conference applications 310A and 310B render three three-dimensional modeling data, including data describing the three-dimensional environment and data representing each participant avatar. This rendering may involve rasterization techniques, texture mapping techniques, ray tracing techniques, shading techniques, or other rendering techniques. In one embodiment, the rendering may involve ray tracing based on the characteristics of the virtual camera. Ray tracing includes tracing the optical path as pixels on the image plane and generating an image by simulating the effect of facing a virtual object. In some embodiments, to enhance realism, ray tracing may simulate optical effects such as reflection, refraction, scattering, and diffusion.

[0053]

[0067] In this way, the user can enter the virtual space using web browsers 308A and 308B. The scene is displayed on the user's screen. The user's web camera video stream and microphone audio stream are sent to the server 302. When other users enter the virtual space, avatar models of those users are created. The position of this avatar is sent to the server and received by other users. Other users also obtain a notification from the server 302 that the audio / video stream is available. The user's video stream is placed on the avatar created for that user. The audio stream is played as coming from the position of the avatar.

[0054]

[0068] Figures 4A - 4C show how data is transferred between various components of the system in FIG. 3 to provide a video conference. As in FIG. 3, each of FIGS. 4A - 4C shows the connection between the server 302 and the devices 306A and 306B. In particular, FIGS. 4A - 4C show an example of the data flow between those devices.

[0055]

[0069] Figure 4A shows a diagram 400 indicating how the server 302 transmits data describing a virtual environment to devices 306A and 306B. In particular, both devices 306A and 306B receive a three-dimensional arena 404, a background texture 402, a spatial hierarchy 408, and any other three-dimensional modeling information 406 from the server 302.

[0056]

[0070] As described above, the background texture 402 is an image showing distinct features of the virtual environment. The image can be regular (such as a brick wall) or irregular. The background texture 402 can be encoded in any common image file format such as a bitmap, JPEG, GIF, or other file image format. For example, describe a background image rendered for a distant sphere.

[0057]

[0071] The three-dimensional arena 404 is a three-dimensional model of the space in which the meeting is held. As described above, the three-dimensional arena 404 can include, for example, a mesh and its own texture information that is probably mapped to the three-dimensional primitives to be described. The virtual camera and each avatar can define the space in which they can navigate within the virtual environment. Thus, it can be delimited by an edge (such as a wall or fence) that shows the user the outer perimeter of the navigable virtual environment.

[0058]

[0072] The spatial hierarchy 408 is data that designates compartments in the virtual environment. These participants are used to specify how the audio is processed before being transferred between participants. As described below, this compartment data can have a hierarchy and can describe the audio processing allowed in areas where participants in the virtual meeting can have private conversations or small talk.

[0059]

[0073] The three-dimensional model 406 is any other three-dimensional modeling information necessary for conducting the meeting. In one embodiment, this can include information describing each avatar. Alternatively or additionally, this information can include a product demo.

[0060]

[0074] With the information necessary to hold the meeting having been sent to the participants, FIGS. 4B and 4C show how the server 302 transfers information between devices. FIG. 4B shows a diagram 420 of how the server 302 receives information from each of the devices 306A and 306B, and FIG. 4C shows a diagram 420 of how the server 302 sends information to each of the devices 306B and 306A. In particular, device 306A sends a position and orientation 422A, a video stream 424A, and an audio stream 426A to the server 302, and the server 302 sends the position and orientation 422A, the video stream 424A, and the audio stream 426A to device 306B. Then device 306B sends a position and orientation 422B, a video stream 424B, and an audio stream 426B to the server 302, and the server 302 sends the position and orientation 422B, the video stream 424B, and the audio stream 426B to device 306A.

[0061]

[0075] The positions and orientations 422A and 422B describe the position and orientation of the virtual camera of the user using device 306A. As described above, the position can be coordinates in three-dimensional space (e.g., x, y, z coordinates), and the orientation can be a direction in three-dimensional space (e.g., pan, tilt, roll). In some embodiments, the user may not be able to control the roll of the virtual camera, and thus the orientation may specify only the pan angle and the tilt angle. Similarly, in some embodiments, the user may not be able to control the z coordinate of the avatar (since the avatar is constrained by virtual gravity), and thus the z coordinate may be unnecessary. In this way, the positions and orientations 422A and 422B may each include at least coordinates on a horizontal plane in three-dimensional virtual space as well as pan values and tilt values. Alternatively or additionally, the user may be able to "jump" the avatar, and thus the Z position may be specified only by an indication of whether the user is jumping the avatar.

[0062]

[0076] In different examples, the positions and orientations 422A and 422B can be transmitted and received using HTTP request responses or using socket messaging.

[0063]

[0077] Video streams 424A and 424B are video data captured from the cameras of respective devices 306A and 306B. The video can be compressed. For example, the video can use any commonly known video codec, including MPEG-4, VP8, or H.264. The video can be captured and transmitted in real time.

[0064]

[0078] Similarly, audio streams 426A and 426B are audio data captured from the microphones of respective devices. The audio can be compressed. For example, the audio can use any commonly known audio codec, including MPEG-4 or Vorbis. The audio can be captured and transmitted in real time. Video stream 424A and audio stream 426A are captured, transmitted, and presented in synchronization with each other. Similarly, video stream 424B and audio stream 426B are also captured, transmitted, and presented in synchronization with each other.

[0065]

[0079] Video streams 424A and 424B and audio streams 426A and 426B can be transmitted using the WebRTC application programming interface. WebRTC is an API available in JavaScript. As described above, devices 306A and 306B download and execute a web application as conference applications 310A and 310B, and conference applications 310A and 310B can be implemented in JavaScript. Conference applications 310A and 310B can receive and transmit video streams 424A and 424B and audio streams 426A and 426B by making API calls from JavaScript using WebRTC.

[0066]

[0080] As mentioned earlier, when a user exits a virtual meeting, this departure is communicated to all other users. For example, when device 306A exits a virtual meeting, server 302 communicates that departure to device 306B. As a result, device 306B stops rendering the avatar corresponding to device 306A and deletes that avatar from the virtual space. Further, device 306B stops receiving video stream 424A and audio stream 426A.

[0067]

[0081] As described above, meeting applications 310A and 310B may periodically or intermittently re-render the virtual space based on new information from each of video streams 424A and 424B, positions and orientations 422A and 422B, and new information related to the three-dimensional environment. For the sake of brevity, each of these updates is described here from the perspective of device 306A. However, those skilled in the art will understand that given similar changes, device 306B will behave similarly.

[0068]

[0082] When device 306A receives video stream 424B, it texture maps the frames from video stream 424A onto the avatar corresponding to device 306B. The texture-mapped avatar is re-rendered within the three-dimensional virtual space and presented to the user of device 306A.

[0069]

[0083] When device 306A receives the new position and orientation 422B, it generates an avatar corresponding to device 306B that is located at the new position and oriented in the new direction. The generated avatar is re-rendered within the three-dimensional virtual space and presented to the user of device 306A.

[0070]

[0084] In some embodiments, server 302 may transmit updated model information that describes a three-dimensional virtual environment. For example, server 302 may transmit updated information 402, 404, 406, or 408. When this occurs, device 306A re-renders the virtual environment based on the updated information. This can be useful when the environment changes over time. For example, an outdoor event may change from midday to dusk as the event progresses.

[0071]

[0085] Here too, when device 306B exits the virtual conference, server 302 sends a notification to device 306A indicating that device 306B is no longer participating in the conference. In that case, device 306A re-renders the virtual environment without the avatar of device 306B.

[0072]

[0086] Although FIGS. 3 in FIGS. 4A - 4C are shown with two devices for simplicity, those skilled in the art will understand that the techniques described herein can be extended to any number of devices. Also, although FIGS. 3 in FIGS. 4A - 4C show a single server 302, those skilled in the art will understand that the functions of server 302 can be distributed across multiple computing devices. In one embodiment, the data transferred in FIG. 4A may be from one network address of server 302, but the data transferred in FIGS. 4B and 4C can be transferred to / from another network address of server 302.

[0073]

[0087] In one embodiment, before entering a virtual meeting, a participant can set up a webcam, microphone, speaker, and graphical settings. In an alternative embodiment, after starting the application, the user can enter a virtual lobby and be greeted in the virtual lobby by an avatar controlled by an actual person. This person can view and change the user's webcam, microphone, speaker, and graphical settings. An attendant can also instruct the user on how to use the virtual environment, for example, by showing, moving around, and interacting. When the user is ready, they automatically exit the virtual waiting room and enter the actual virtual environment.

[0074] Volume adjustment in a video conference in a virtual environment

[0088] The embodiment also adjusts the volume to provide a sense of position and space within the virtual meeting. This is shown, for example, in FIGS. 5-7, FIGS. 8A, 8B, and FIGS. 9A-9C, and will be described below for each figure.

[0075]

[0089] FIG. 5 is a flowchart showing a method 500 for adjusting the relative left and right volumes to provide a sense of position in a virtual environment during a video conference.

[0076]

[0090] In step 502, the volume is adjusted based on the distance between the avatars. As described above, an audio stream from the microphone of another user's device is received. The volume of both the first and second audio streams is adjusted based on the distance between the first position and the second position. This is shown in FIG. 6.

[0077]

[0091] Figure 6 shows a chart 600 indicating how the volume decreases as the distance between avatars increases. Chart 600 shows the volume 602 on the x-axis and y-axis. As the distance between users increases, the volume remains constant until the reference distance 602 is reached. When the reference distance 602 is reached, the volume begins to decrease. In this way, all other things being equal, closer users will have a louder sound than farther users.

[0078]

[0092] The rate at which the voice decreases depends on the attenuation coefficient. This can be a coefficient built into the settings of the video conferencing system or the client device. As shown by line 608 and line 610, the greater the attenuation coefficient, the more rapidly the volume decreases compared to a smaller one.

[0079]

[0093] Returning to Figure 5, in step 504, the relative left and right audio is adjusted based on the direction in which the avatars are placed. That is, the volume of the audio output by the user's speaker (e.g., headset) is changed to provide a sense of the location where the avatar of the speaking user is placed. The relative volume of the left and right audio streams is adjusted based on the direction of the location where the user generating the audio stream is placed (e.g., the location of the avatar of the speaking user) relative to the location where the user receiving the audio is placed (e.g., the location of the virtual camera). The location can be on a horizontal plane within the three-dimensional virtual space. The relative volume of the left and right audio is streamed so as to provide a sense of the location where the second position exists relative to the first position in the three-dimensional virtual space.

\(0080\)

[0094] For example, in step 504, the audio corresponding to the avatar on the left side of the virtual camera is adjusted so that the audio is output at a louder volume in the left ear than in the right ear of the receiving user. Similarly, the audio corresponding to the avatar on the right side of the virtual camera is adjusted so that the audio is output at a louder volume in the right ear than in the left ear of the receiving user.

\(0081\)

[0095] In step 506, the relative left and right audio is adjusted based on the direction in which one avatar is facing the other avatar. The relative volume of the left and right audio streams is adjusted based on the angle between the direction the virtual camera is facing and the direction the avatar is facing, such that the volume difference between the left and right audio streams tends to increase as the angle becomes perpendicular.

[0082]

[0096] For example, when the avatar is facing the virtual camera directly, the relative left and right volumes of the avatar corresponding to the audio stream may not be adjusted at all in step 506. When the avatar is facing the left side of the virtual camera, the relative left and right volumes of the avatar corresponding to the audio stream can be adjusted such that the left is louder than the right. And when the avatar is facing the right side of the virtual camera, the relative left and right volumes of the avatar corresponding to the audio stream can be adjusted such that the right is louder than the left.

[0083]

[0097] In one example, the calculation in step 506 can include taking the cross product of the angle the virtual camera is facing and the angle the avatar is facing. The angles can be the directions facing on the horizontal plane.

[0084]

[0098] In one embodiment, a check can be performed to identify the audio output device the user is using. If the audio output device is not a set of headphones or another type of speaker that provides stereo effects, the adjustments in steps 504 and 506 may not be performed.

[0085]

[0099] Steps 502 - 506 are repeated for every audio stream received from any other participant. Based on the calculations in steps 502 - 506, the left and right audio gains for any other participant are calculated.

[0086]

[0100] In this way, the audio stream of each participant is adjusted to provide a sense of the location where the participant's avatar is placed in the three-dimensional virtual environment.

[0087]

[0101] Not only is the audio stream adjusted to provide a sense of the location where the avatar is placed, but in certain embodiments, the audio stream can also be adjusted to provide a private or semi-private volume area. In this way, the virtual environment enables the user to have a private conversation. Also, the virtual environment enables the users to communicate with each other and have individual small talks that were not possible with conventional video conferencing software. This is shown, for example, with respect to FIG. 7.

[0088]

[0102] FIG. 7 is a flowchart showing a method 700 of adjusting relative volumes to provide different volume areas in a virtual environment during a video conference.

[0089]

[0103] As described above, the server may provide the specifications of the voice or volume area to the client device. The virtual environment may be partitioned into different volume areas. In step 702, the device identifies in which voice area each avatar and the virtual camera are placed.

[0090]

[0104] For example, FIGS. 8A and 8B are diagrams showing different volume areas in a virtual environment during a video conference. FIG. 8A shows an interface 800 having a volume area 802 that enables a semi-private conversation or small talk between the user controlling avatar 806 and the user controlling the virtual camera. In this way, the users around the conference table 810 can talk without disturbing the others in the room. The voice from the user controlling avatar 806 of the virtual camera may decrease as it exits volume area 802, but it does not completely disappear. Thereby, passersby can join the conference if they want to participate.

[0091]

[0105] Interface 800 also includes buttons 804, 806, and 808 described below.

[0092]

[0106] Figure 8B shows FIG. 800 having a volume area 804 that enables a private conversation between the user controlling avatar 808 and the user controlling the virtual camera. When entering the volume area 804, only the audio from the user controlling avatar 808 and the user controlling the virtual camera can be output to the user inside the volume area 804. If the audio is not played at all from those users to other users in the meeting, those audio streams may not even be sent to other user devices.

[0093]

[0107] The volume space can be hierarchical as shown in FIGS. 9A and 9B. FIG. 9B shows a layout having different volume areas arranged in a hierarchy. Volume areas 934 and 935 are within volume area 933, and volume areas 933 and 932 are within volume area 931. These volume areas are represented by a hierarchical tree as shown in FIGS. 900 and 9A.

[0094]

[0108] In FIG. 900, node 901 represents volume area 931 and is the root of the tree. Nodes 902 and 903 are children of node 901 and represent volume areas 932 and 933. Nodes 904 and 906 are children of node 903 and represent volume areas 934 and 935.

[0095]

[0109] If a user located in area 934 tries to hear a speaking user located in area 932, the audio stream needs to pass through several different virtual "walls" that each attenuate the audio stream. In particular, the sound needs to pass through the wall of area 932, the wall of area 933, and the wall of area 934. Each wall attenuates by a specific factor. This calculation is described with respect to steps 704 and 706 of FIG. 7.

[0096]

[0110] In step 704, traverse the hierarchy to identify various audio areas between the avatars. This is shown, for example, in FIG. 9C. Starting from the node corresponding to the virtual area of the voice (in this case, node 904), the path to the node of the receiving user (in this case, node 902) is identified. To identify the path, the link 952 between the nodes is identified. In this way, a subset of the areas between the area containing the avatar and the area containing the virtual camera is identified.

[0097]

[0111] In step 706, the audio stream from the speaking user is attenuated based on the wall transmission coefficient of each of the subset of areas. Each wall transmission coefficient specifies the amount by which the audio stream is attenuated.

[0098]

[0112] Additionally or alternatively, different areas may in that case have different attenuation factors, and the distance-based calculations shown in method 600 may be applied to the individual areas based on each attenuation factor. In this way, different areas of the virtual environment sound at different rates. The audio gain identified in the method described above with respect to FIG. 5 may be applied to the audio stream, and the left and right audio may be identified accordingly. In this way, both the wall transmission coefficient and the attenuation factor, as well as the left-right adjustment, for providing a directional separation in the audio may be applied together to provide a comprehensive audio experience.

[0099]

[0113] Different audio areas may have different functions. For example, the volume area may be a podium area. If the user is located in the podium area, some or all of the attenuation described with respect to FIG. 5 or FIG. 7 may not occur. For example, attenuation may not occur due to the attenuation factor or the wall transmission coefficient. In some embodiments, the relative left and right audio may still be adjusted to provide a sense of direction.

[0100]

[0114] For illustrative purposes, the method described with respect to FIGS. 5 and 7 describes an audio stream from a user having a corresponding avatar. However, the same method can be applied to other sound sources other than avatars. For example, the virtual environment can have a three-dimensional model of a speaker. For a presentation or simply to provide background music, sound can be emitted from the speaker in the same manner as the avatar model described above.

[0101]

[0115] As mentioned above, the wall transmission coefficient can be used to globally separate the audio. In one embodiment, this can be used to create a virtual office. In one example, each user can have a monitor in their physical (perhaps home) office that is always on and displays a conferencing application logged into the virtual office. There may be a feature that allows the user to indicate whether they are in the office and not available for entry. When the unavailable indicator is off, a colleague or manager can casually visit within the virtual space and knock or enter in the same way as a physical office. A visitor may be able to leave a memo if the operator is not in the office. When the operator returns, the operator can read the memo left by the visitor. The virtual office can have a whiteboard and / or interface that displays messages to the user. The messages may be emails and / or from a messaging application such as the SLACK application available from Slack Technologies, Inc. located in San Francisco, CA.

[0102]

[0116] The user may be able to customize or personalize their virtual office. For example, the user may be able to adorn a model of a poster or other wall decoration. The user may be able to change the model or orientation of a decorative object such as a desk or plant. The user may be able to change the lighting or the view from the window.

[0103]

[0117] Returning to FIG. 8A, the interface 800 includes various buttons 804, 806, and 808. When the user presses button 804, the attenuation described above with respect to the methods of FIGS. 5 and 7 can be performed, or can be performed only in smaller amounts. In that situation, the user's voice is output uniformly to other users, enabling the user to provide their speech to all participants in the meeting. The user video can also be output to the presentation screen within the virtual environment in the same manner as described below. When the user presses button 806, the speaker mode is enabled. In that case, the audio is output from the sound source within the virtual environment for playing background music or the like. When the user presses button 808, screen sharing can be enabled, allowing the user to share the content of the screen or window on their respective device with other users. The content can be presented in the presentation model. This will also be described below.

[0104] Presentation in a Three-Dimensional Environment

[0118] FIG. 10 shows an interface 1000 having a three-dimensional model 1004 in a three-dimensional virtual environment. As described above with respect to FIG. 1, the interface 1000 can be displayed to a user who can navigate around the virtual environment. As shown in the interface 1000, the virtual environment includes an avatar 1004 and a three-dimensional model 1002.

[0105]

[0119] The three-dimensional model 1002 is a 3D model of a product placed inside the virtual space. People can join this virtual space and can observe the model and walk around it. The product can have local audio to enhance the experience.

[0106]

[0120] More specifically, when a presenter in a virtual space wants to show a 3D model, the user selects the desired model from the interface. This sends a message to the server that updates the details (including the name and path of the model). This is automatically communicated to the client. In this way, the three-dimensional model can be rendered to be displayed simultaneously with the presentation of the video stream. The user can navigate a virtual camera around the three-dimensional model of the product.

[0107]

[0121] In different examples, the object may be a product demo or an advertisement for a product.

[0108]

[0122] FIG. 11 shows an interface 1100 having presentation screen sharing in a three-dimensional virtual environment used in a video conference. As described above with respect to FIG. 1, the interface 1100 can be displayed to a user who can navigate around the virtual environment. As shown in the interface 1100, the virtual environment includes an avatar 1104 and a presentation screen 1106.

[0109]

[0123] In this embodiment, a presentation stream from the devices of the participants in the meeting is received. The presentation stream is texture mapped to the three-dimensional model of the presentation screen 1106. In one embodiment, the presentation stream can be a video stream from the camera of the user's device. In another embodiment, the presentation stream can be screen sharing from the user's device, in which case a monitor or window is shared. Through screen sharing or other means, the presentation video and audio stream can also be from an external source, such as a live stream of an event. When the user enables the presentation mode, the user's presentation stream (and audio stream) is tagged with the name of the screen the user wants to use and published to the server. Other clients are notified that a new stream is available.

[0110]

[0124] The presenter may also be able to control the location and orientation of the audience members. For example, the presenter may have an option to be selected to relocate all other participants to the meeting so that they are positioned and oriented to face the presentation screen.

[0111]

[0125] The audio stream is captured from the microphone of the first participant's device in synchronization with the presentation stream. The audio stream from the user's microphone can be heard by other users as if it were from the presentation screen 1106. In this way, the presentation screen 1106 can be a sound source as described above. Since the user's audio stream is projected from the presentation screen 1106, the one from the user's avatar can be suppressed. In this way, the audio stream is output and played back in synchronization with the display of the presentation stream on the screen 1106 in the three-dimensional virtual space.

[0112] Bandwidth allocation based on the distance between users

[0126] FIG. 12 is a flowchart showing a method 1200 for allocating available bandwidth based on the relative positions of avatars in a three-dimensional virtual environment.

[0113]

[0127] In step 1202, the distance between a first user and a second user in the virtual meeting space is determined. The distance can be the distance between the users on the horizontal plane in the three-dimensional space.

[0114]

[0128] In step 1204, the received video stream is prioritized such that the video stream from a nearer user is given a higher priority than the video stream from a farther user. The priority value can be determined as shown in FIG. 13.

[0115]

[0129] Figure 13 shows a chart 1300 indicating the priority 1306 and distance 1302 on the y-axis. As shown by line 1306, the priority state remains constant until the reference distance 1304 is reached. After reaching the reference distance, the priority begins to decrease.

[0116]

[0130] In step 1206, the available bandwidth to the user device is distributed among various video streams. This can be done based on the priority values identified in step 1204. For example, the priorities can be proportionally adjusted so that they all add up to 1 when combined. For any video with insufficient available bandwidth, the relative priority can be set to zero. Then, the priorities are adjusted again for the remaining video streams. The bandwidth is allocated based on these relative priority values. Additionally, bandwidth can be reserved for audio streams. This is shown in Figure 14.

[0117]

[0131] Figure 14 shows a chart 1400 having a y-axis representing the bandwidth 1406 and an x-axis representing the relative priority. After the minimum bandwidth 1406 required for a video to be valid is allocated to the video, the bandwidth 1406 allocated to the video stream increases in proportion to its relative priority.

[0118]

[0132] Once the allocated bandwidth is determined, the client can request the video from the server at the bandwidth / bitrate / frame rate / resolution selected and allocated for that video. This can start a negotiation process between the client and the server to start streaming the video at the specified bandwidth. In this way, the available video and audio bandwidth is properly divided among all users, and a user with twice the priority will obtain twice as much bandwidth.

[0119]

[0133] In one possible implementation, using simulcast, all clients send multiple video streams to the server at different bitrates and resolutions. Other clients can then indicate to the server one of these streams that the client is interested in receiving.

[0120]

[0134] In step 1208, it is determined whether the available bandwidth between the first user and the second user in the virtual conference space is such that the display of the video at that distance is inefficient. This determination can be made by either the client or the server. When made by the client, the client sends a message to the server to stop video transmission to the client. If it is inefficient, the transmission of the video stream to the second user's device is stopped, and the second user's device is notified to replace the video stream with a still image. The still image can simply be the last video frame received (or one of the last video frames).

[0121]

[0135] In one embodiment, a similar process can be performed for audio, and given the size of the portion reserved for audio, the quality can be reduced. In another embodiment, a fixed bandwidth is given to each audio stream.

[0122]

[0136] In this way, the embodiment can improve the performance of all users and the server, and reduce the quality of video streams and audio streams for users who are far away and / or have low importance. This is not done when a sufficient bandwidth budget is available. The reduction is done in both bitrate and resolution. Since the encoder can use the available bandwidth for that user more efficiently, this improves the quality of the video.

[0123]

[0137] Independently of this, the video resolution is reduced based on the distance, and a user who is twice as far has half the resolution. In this way, unnecessary resolution does not need to be downloaded given the limit of the screen resolution. Thus, bandwidth is conserved.

[0124]

[0138] FIG. 15 is a diagram of a system 1500 showing the components of a device used to provide a video conference within a virtual environment. In various embodiments, system 1500 can operate according to the methods described above.

[0125]

[0139] Device 306A is a user computing device. Device 306A can be a desktop or laptop computer, a smartphone, a tablet, or a wearable (such as a watch or a head-mounted display). Device 306A includes a microphone 1502, a camera 1504, a stereo speaker 1506, and an input device 1512. Although not shown, device 306A also includes a processor and a persistent non-volatile memory. The processor can include one or more central processing units, a graphics processing unit, or any combination thereof.

[0126]

[0140] Microphone 1502 converts sound into an electrical signal. Microphone 1502 is positioned to capture the speech of the user of device 306A. In different examples, microphone 1502 can be a condenser microphone, an electret microphone, a moving coil microphone, a ribbon microphone, a carbon microphone, a piezoelectric microphone, an optical fiber microphone, a laser microphone, a water microphone, or a MEMS microphone.

[0127]

[0141] Camera 1504 captures image data by generally capturing light through one or more lenses. Camera 1504 is positioned to capture a photographic image of a user of device 306A. Camera 1504 includes an image sensor (not shown). The image sensor can be, for example, a charge-coupled device (CCD) sensor or a complementary metal-oxide semiconductor (CMOS) sensor. The image sensor can include one or more photodetectors that detect light and convert it into an electrical signal. These electrical signals captured together within a similar time frame constitute a still photographic image. A series of still photographic images captured together at regular intervals constitutes a video. In this way, camera 1504 captures images and videos.

[0128]

[0142] Stereo speaker 1506 is a device that converts an electrical audio signal into corresponding left and right sounds. Stereo speaker 1506 outputs a left audio stream and a right audio stream that are generated by an audio processor 1520 (hereinafter) and played back stereophonically to a user of device 306A. Stereo speaker 1506 includes both ambient speakers and headphones designed to play sound directly into the user's left and right ears. Examples of speakers include moving iron loudspeakers, piezoelectric speakers, electrostatic loudspeakers, electrostatic loudspeakers, ribbons and planar magnetic loudspeakers, bending wave speakers, flat panel speakers, high-level air motion transducers, transparent ion conducting speakers, plasma arc speakers, thermoacoustic speakers, rotary woofers, moving coils, electrostatic, electret, planar magnetic, and balanced armchairs.

[0129]

[0143] Network interface 1508 is a software or hardware interface between two devices or between two protocol layers within a computer network. Network interface 1508 receives the video streams of each participant in the meeting from server 302. The video streams are captured from the cameras of the devices of the other participants in the videoconference. Network interface 1508 also receives data specifying the three-dimensional virtual space and any models within it from server 302. For each of the other participants, network interface 1508 receives the position and orientation in the three-dimensional virtual space. The position and orientation are input by each of the other participants.

[0130]

[0144] Network interface 1508 also sends data to server 302. Network interface 1508 sends the position of the virtual camera of the user of device 306A used by renderer 1518, and sends video and audio streams from camera 1504 and microphone 1502.

[0131]

[0145] Display 1510 is an output device for presenting electronic information in visual or tactile form (the tactile form is used, for example, in tactile electronic displays for visually impaired people). Display 1510 can be a television set, a computer monitor, a head-mounted display, a head-up display, the output of an augmented reality or virtual reality headset, a broadcast reference monitor, a medical monitor, a mobile display (of a mobile device), a smartphone display (of a smartphone). To present information, display 1510 can include an electroluminescent (ELD) display, a liquid crystal display (LCD), a light-emitting diode (LED) backlit LCD, a thin-film transistor (TFT) LCD, a light-emitting diode (LED) display, an OLED display, an AMOLED display, a plasma (PDP) display, a quantum dot (QLED) display.

[0132]

[0146] The input device 1512 is a device used to provide data and control signals to an information processing system such as a computer or an information appliance. The input device 1512 enables a user to input a new desired position of a virtual camera used by the renderer 1518, thereby enabling navigation in a three-dimensional environment. Examples of input devices include keyboards, mice, scanners, joysticks, and touchscreens.

[0133]

[0147] The web browser 308A and the web application 310A have been described above with respect to FIG. 3. The web application 310A includes a screen capture 1514, a texture mapper 1516, a renderer 1518, and an audio processor 1520.

[0134]

[0148] The screen capture 1514 captures a presentation stream, particularly screen sharing. The screen capture 1514 can interact with an API provided by the web browser 308A. By calling functions available from the API, the screen capture 1514 can prompt the web browser 308A to ask the user which window or screen they want to share. Based on the answer to that query, the web browser 308A can return a video stream corresponding to the screen sharing to the screen capture 1514, and the screen capture 1514 passes it to the network interface 1508 for transmission to the server 302 and ultimately to the devices of other participants.

[0135]

[0149] The texture mapper 1516 texture maps a video stream onto a three-dimensional model corresponding to an avatar. The texture mapper 1516 can texture map each frame from the video onto the avatar. In addition, the texture mapper 1516 can texture map a presentation stream onto a three-dimensional model of a presentation screen.

[0136]

[0150] The renderer 1518 renders a three-dimensional virtual space that includes texture-mapped three-dimensional models of each participant's avatar positioned and oriented in corresponding received positions for output to the display 1510 from the perspective of the virtual camera of the user of the device 306A. The renderer 1518 also renders any other three-dimensional models, such as for example a presentation screen.

[0137]

[0151] The audio processor 1520 adjusts the volume of the received audio stream to identify a left audio stream and a right audio stream and to provide a sense of where the second position is in the three-dimensional virtual space relative to the first position. In one embodiment, the audio processor 1520 adjusts the volume based on the distance between the first position and the second position. In another embodiment, the audio processor 1520 adjusts the volume based on the direction of the second position relative to the first position. In yet another embodiment, the audio processor 1520 adjusts the volume based on the direction of the second position relative to the first position on a horizontal plane within the three-dimensional virtual space. In yet another embodiment, the audio processor 1520 adjusts the volume based on the direction that the virtual camera faces within the three-dimensional virtual space such that when the avatar is positioned to the left of the virtual camera, the left audio stream tends to have a greater volume, and when the avatar is positioned to the right of the virtual camera, the right audio stream tends to have a greater volume. Finally, in yet another embodiment, the audio processor 1520 adjusts the volume based on the angle between the direction that the virtual camera faces and the direction that the avatar faces such that the closer the angle is to being perpendicular to the direction that the avatar faces, the greater the difference in volume between the left and right audio streams tends to be.

[0138]

[0152] The audio processor 1520 can also adjust the volume of the audio stream based on the area where the speaker is located with respect to the area where the virtual camera is located. In this embodiment, the three-dimensional virtual space is segmented into a plurality of areas. These areas can have a hierarchy. When the speaker and the virtual camera are arranged in different areas, a wall transmission coefficient can be applied to attenuate the volume of the speech audio stream.

[0139]

[0153] The server 302 includes an attendee notifier 1522, a stream adjuster 1524, and a stream transmitter 1526.

[0140]

[0154] The attendee notifier 1522 notifies the meeting participants when a participant joins and leaves the meeting. When a new participant joins the meeting, the attendee notifier 1522 sends a message indicating that a new participant has joined to the devices of the other participants in the meeting. The attendee notifier 1522 signals the stream transmitter 1526 to start forwarding video, audio, and location / direction information to the other participants.

[0141]

[0155] The stream adjuster 1524 receives a video stream captured from the camera of the first user's device. The stream adjuster 1524 identifies the bandwidth available for transmitting the virtual meeting data to the second user. The stream adjuster 1524 identifies the distance between the first user and the second user in the virtual meeting space. Then the stream adjuster 1524 distributes the available bandwidth between the first video stream and the second video stream based on the relative distance. In this way, the stream adjuster 1524 assigns a higher priority to the video stream of the nearer user than to the video stream of the farther user. Additionally or alternatively, the stream adjuster 1524 may be arranged in the device 306A, perhaps as part of the web application 310A.

[0142]

[0156] The stream transmitter 1526 broadcasts the received position / direction information, video, audio, and screen sharing screen (in a state adjusted by the stream adjuster 1524). The stream transmitter 1526 can transmit information to the device 306A in response to a request from the conference application 310A. The conference application 310A can transmit the request in response to a notification from the attendee notifier 1522.

[0143]

[0157] The network interface 1528 is a software or hardware interface between two devices or between two protocol layers within a computer network. The network interface 1528 transmits model information to the devices of various participants. The network interface 1528 receives video, audio, and screen sharing screens from various participants.

[0144]

[0158] The screen capture 1514, texture mapper 1516, renderer 1518, audio processor 1520, attendee notifier 1522, stream adjuster 1524, and stream transmitter 1526 can each be implemented in hardware, software, firmware, or any combination thereof.

[0145]

[0159] Identifiers such as "(a)", "(b)", "(i)", "(ii)", etc. may be used for different elements or steps. These identifiers are used for clarity and do not necessarily indicate the order of the elements or steps.

[0146]

[0160] The present invention has been described above using functional building blocks that indicate the implementation of the specified functions and their relationships. The boundaries of these functional building blocks are arbitrarily defined in this specification for convenience of explanation. Alternative boundaries can be defined as long as the specified functions and their relationships are appropriately executed.

[0147]

[0161] The foregoing description of the specific embodiments will enable others skilled in the art to make various changes and / or adaptations to the various uses, such as the specific embodiments, without undue experimentation and without departing from the general concept of the invention. Thus, such adaptations and changes are intended to be within the meaning and scope of the disclosed embodiments and equivalents thereof, based on the teachings and guidance presented herein. It should be understood that the expressions and terms herein are for the purpose of description and not limitation, and thus the terms or expressions herein should be interpreted by those skilled in the art in light of the teachings and guidance.

[0148]

[0162] The breadth and scope of the present invention should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

Claims

1. A computer-implemented method for streaming a video of a virtual meeting, comprising: (a) determining a distance between a first user and a second user in a virtual meeting space; (b) receiving a video stream captured from a first camera of a first device of the first user, wherein the first camera is positioned to capture a photographic image of the first user; (c) based on the determined distance, selecting a reduced resolution or bitrate of the video stream such that the closer the distance, the higher the resolution or bitrate compared to when the distance is far; (d) requesting transmission of the video stream at the reduced resolution or bitrate to a second device of the second user for display within the virtual meeting space, wherein the virtual meeting space is rendered by a web conferencing application running on a web browser on the second device of the second user, and the video stream is mapped to an avatar of the first user for display to the second user within the virtual meeting space; (e) determining that the distance between the first user and the second user in the virtual meeting space is such that video display at that distance is inefficient; in response to the determination in (e), (f) stopping transmission of the video stream to the second device of the second user; (g) notifying the second device of the second user to replace the video stream with a still image; A method comprising the above steps.

2. (h) receiving a second video stream captured from a third camera of a third device of a third user, wherein the third camera is positioned to capture a photographic image of the third user; (i) determining a bandwidth available for transmitting virtual meeting data to the second user; (j) determining a second distance between the third user and the second user in the virtual meeting space; Based on the second distance specified in (j) relative to the distance specified in (k)(a), distribute the available bandwidth between the video stream received in (b) and the second video stream received in (h). The method according to claim 1, further comprising.

3. The distributing (k) includes assigning a higher priority to the video stream of a user closer than the video stream from a distant user, the method according to claim 2.

4. (l) Receiving a first audio stream from the first device of the first user. (m) Receiving a second audio stream from the third device of the third user, wherein the distributing (k) includes receiving, including securing a portion of the first and second audio streams. The method according to claim 2, further comprising.

5. The method according to claim 4, further comprising (n) reducing the quality of the first and second audio streams according to the size of the secured portion.

6. The reducing (n) includes reducing the quality independently of the second distance specified in (j) relative to the distance specified in (a), the method according to claim 5.

7. The video stream at the reduced resolution is mapped to the avatar rendered at the position of the first user in the virtual meeting space by the second device of the second user for display to the second user, the method according to claim 1.

8. A non-transitory tangible computer-readable device storing instructions that, when executed by at least one computing device, cause the at least one computing device to perform operations for streaming video of a virtual meeting, the operations including (a) Identifying the distance between a first user and a second user in a virtual meeting space. (b) Receiving a video stream captured by a first camera of a first device of the first user, wherein the first camera is positioned to capture a photographic image of the first user, receiving. (c)Based on the specified distance, select a reduced resolution or bit rate of the video stream such that the closer the distance, the higher the resolution or bit rate compared to when the distance is far. (d)Request the transmission of the video stream at the reduced resolution or bit rate to the second device of the second user for display within the virtual meeting space, where the virtual meeting space is rendered by a web conferencing application running on a web browser on the second device of the second user, and the video stream is mapped to the avatar of the first user for display to the second user within the virtual meeting space. (e)Specify that the distance between the first user and the second user in the virtual meeting space is such that the display of video at that distance is inefficient. In response to the specification in (e) (f)Abort the transmission of the video stream to the second device of the second user. (g)Notify the second device of the second user to replace the video stream with a still image. A device comprising the above.

9. The operation is (h)Receive a second video stream captured from a third camera of a third device of a third user, where the third camera is positioned to capture a photographic image of the third user. (i)Specify the bandwidth available for transmitting the virtual meeting data to the second user. (j)Specify a second distance between the third user and the second user in the virtual meeting space. (k)Based on the second distance specified in (j) relative to the distance specified in (a), allocate the available bandwidth between the video stream received in (b) and the second video stream received in (h). The device according to claim 8, further comprising the above.

10. The allocating in (k) includes assigning a higher priority to the video stream of a nearer user than to the video stream of a farther user. The device according to claim 9.

11. The operation is Receiving a first audio stream from a first device of the first user; Receiving a second audio stream from a third device of the third user, wherein the distributing (k) includes securing a portion of the first and second audio streams; The device according to claim 9, further comprising: **Claim 12** The operation further includes: The device according to claim 11, further comprising reducing the quality of the first and second audio streams according to the size of the secured portion. **Claim 13** The device according to claim 12, wherein the reducing (n) includes reducing the quality independently of the second distance specified in (j) relative to the distance specified in (a). **Claim 14** The video stream at the reduced resolution is mapped to an avatar rendered at the position of the first user in the virtual meeting space by the second device of the second user for display to the second user, according to the device of claim 8. **Claim 15** A system for streaming video for a virtual meeting, comprising: A processor coupled to a memory; A network interface for receiving a video stream captured by a first camera of a first device of a first user, the first camera being positioned to capture a photographic image of the first user; A stream adjuster configured to identify a distance between a first user and a second user in a virtual meeting space and, based on the identified distance, reduce the resolution of the video stream such that the closer the distance, the higher the resolution or bit rate compared to when the distance is far; Comprising: The network interface is configured to transmit the video stream at a reduced resolution to a second device of the second user for display to the second user in the virtual meeting space, the virtual meeting space is rendered by a web conferencing application running on a web browser on the second device of the second user, and the video stream is mapped to an avatar of the first user for display to the second user in the virtual meeting space. The stream adjuster is further configured to, in response to identifying that the distance between the first user and the second user in the virtual conference space is such that the display of the video at the distance is inefficient, stop transmitting the video stream to the second device of the second user and notify the second device of the second user to replace the video stream with a still image. System. **Claim 16** The network interface is configured to receive a second video stream captured from a third camera of a third device of a third user, and the third camera is positioned to capture a photographic image of the third user. The system according to claim 15, wherein the stream adjuster (i) identifies the bandwidth available for transmitting the virtual conference data to the second user, (ii) identifies a second distance between the third user and the second user in the virtual conference space, and (iii) distributes the available bandwidth to the video stream and the second video stream based on the second distance relative to the distance. **Claim 17** The system according to claim 16, wherein the stream adjuster is configured to assign a higher priority to the video stream of a user closer than the video stream from a distant user. **Claim 18** The system according to claim 16, wherein the network interface is configured to receive a first audio stream from a first device of the first user and a second audio stream from a third device of the third user, and the stream adjuster is configured to reserve a portion of the first and second audio streams.

Citation Information

Patent Citations

  • Roughing of resin layer securely without corrosion and cracking

    JP1986090497A

  • Three-dimensional virtual space sharing device

    JP1995288791A

  • Method and device for three-dimensional virtual space sound communication

    JP1998055261A

  • Virtual space sharing system and terminal equipment and repeater system and server and computer readable record medium for recording program for terminal equipment and child information relay program

    JP1999175450A

  • Video transmission method, intermediation server device and program recording medium

    JP2001094963A