Live virtual remote in a browser

By segmenting and projecting the images of video call participants in the browser, the problem of social distance in video calls is solved, enabling advanced video calling features with zero installation, improving user experience and reducing resource consumption.

CN115968544BActive Publication Date: 2026-05-08GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GOOGLE LLC
Filing Date
2020-08-24
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In video calls, participants are presented in their own environment, leading to a sense of social distance, and existing technologies require the installation of full-featured applications on user devices to achieve advanced features such as background modification.

Method used

By implementing virtual remote video calls in a browser, using machine learning models to segment video frames and select participant image segments, the video is directly streamed to another device and projected and rendered on the receiving device, achieving a zero-installation application environment.

Benefits of technology

It enables real-time virtual remote transmission of participant images without installing additional applications, enhancing the social experience, reducing resource consumption, and providing flexible video calling capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115968544B_ABST
    Figure CN115968544B_ABST
Patent Text Reader

Abstract

A method comprising opening a web-based video call in a browser on a first device (145), receiving a request to join the web-based video call from a second device (150), capturing (110) video comprising frames (105) by the first device, segmenting (115) the frames by the first device, selecting at least one segment (120) of the segmented frames by the first device, and streaming (125) the video comprising the at least one segment as a live virtual teleport (140) directly from the first device to the second device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The examples relate to streaming video in a web-based video conferencing environment. Background Technology

[0002] Video calls can make users feel disconnected from each other. In other words, social interaction can feel distant because two or more participants are in different locations, each viewing a different location or artificial background on a viewing device (e.g., a mobile phone). Furthermore, for video conferencing to have advanced features (e.g., background modification), a fully-featured application needs to be installed on the user's device. Summary of the Invention

[0003] In general, an apparatus, system, non-transitory computer-readable medium (on which computer-executable program code capable of being executed on a computer system) and / or method can perform a process by means of: opening a web-based video call in a browser on a first device; receiving a request to join the web-based video call from a second device; capturing video including frames by the first device; segmenting the frames by the first device; selecting at least one segment of the segmented frames by the first device; and streaming the video including the at least one segment directly from the first device to the second device as a real-time virtual teletransmission.

[0004] The implementation may include one or more of the following features. For example, opening the web-based video call includes loading a webpage containing code configured to implement a trained machine learning model, which can be configured to segment the frame and select the at least one segment. The at least one segment may be an image of a participant in the web-based video call. The web-based video call may be implemented using a web-based communication standard. The segmentation of the frame may include: grouping pixels in the frame into semantic regions to locate objects and boundaries, classifying the pixels of the frame into two categories: 1) pixels representing people, and 2) pixels representing background, and segmenting the pixels representing people from the frame. The segmentation of the frame may include identifying each object in the frame, selecting the at least one segment includes selecting an object as the at least one segment, and the object may be a participant in the web-based video call. The at least one segment may be an image of a participant in the web-based video call, and the method may further include converting the image from a two-dimensional image to a three-dimensional image. The at least one segment may be an image of a participant in the web-based video call, and the method may further include applying a filter to the image. The web-based video call can be implemented with zero web application installation.

[0005] In another general aspect, an apparatus, system, non-transitory computer-readable medium (on which computer-executable program code capable of being executed on a computer system) and / or method can perform a process by means of: opening a web-based video call webpage in a browser on a first device; transmitting a request from the first device to join a web-based video call from a second device; receiving streaming video directly from the second device at the first device as a first video; capturing a second video by the first device; orienting the first video based on the second video by the first device; projecting the first video onto the second video by the first device to generate a third video; and rendering a webpage including the third video by the first device.

[0006] The implementation may include one or more of the following features. For example, the method may further include generating a plane, and the orientation of the first video may include: determining a normal vector associated with the second video, and rotating and translating at least one of the first videos based on the normal vector. The method may further include generating a plane and positioning the plane in the second video, wherein projecting the first video into the second video includes adding the first video to the plane. The method may further include generating a plane and positioning the plane in the second video, and the orientation of the first video may include: determining a normal vector associated with the second video, and rotating and translating at least one of the first videos based on the normal vector, and projecting the first video into the second video may include adding the first video to the plane. The first video may be of a first participant in the web-based video call, and the second video may be a real-world video. The plane may be a transparent two-dimensional virtual structure located in the second video. The plane may have a size proportional to the display of the device rendering the web-based video call webpage. The web-based video call webpage may include code configured to implement a trained machine learning model, and the web-based video call webpage may include code configured to implement web-based augmented reality tools. Web-based video call web pages and the web-based video call itself can be implemented with zero web application installation.

[0007] In yet another general aspect, an apparatus, system, non-transitory computer-readable medium (on which computer-executable program code capable of being executed on a computer system) and / or method can perform a process by means of: opening a web-based video call in a browser on a first device; receiving a request to join the web-based video call from a second device; capturing a first video including frames by the first device; segmenting the frames by the first device; selecting at least one segment of the segmented frames by the first device; streaming the first video including the at least one segment directly from the first device to the second device as a first real-time virtual teleport image; receiving the streamed video directly from the second device by the first device as a second video, the second video including a second real-time virtual teleport image; capturing a third video by the first device; orienting the second video based on the third video by the first device; projecting the second video onto the third video by the first device to generate a fourth video including the second real-time virtual teleport image; and rendering a webpage including the fourth video by the first device.

[0008] The implementation may include one or more of the following features. For example, initiating the web-based video call may include loading a webpage, the webpage may include code configured to implement a trained machine learning model, the trained machine learning model may be configured to segment the frames and select the at least one segment, and the webpage may include code configured to implement a web-based augmented reality tool. The web-based video call may be implemented as a zero-installation web application. The segmentation of the frames may include identifying each object in the frames, selecting the at least one segment may include selecting an object as the at least one segment, and the object may be a participant in the web-based video call. The method may further include generating a plane and positioning the plane in the second video, and the orientation of the first video may include: determining a normal vector associated with the second video, and rotating and translating at least one of the first videos based on the normal vector, and projecting the first video onto the second video may include adding the first video to the plane. Attached Figure Description

[0009] Exemplary embodiments will be more fully understood from the following detailed description and accompanying drawings, wherein like elements are designated by like reference numerals, which are given by way of illustration only and therefore do not limit the exemplary embodiments, and wherein:

[0010] Figure 1 A block diagram illustrating a signal flow according to at least one exemplary embodiment is shown.

[0011] Figure 2A A block diagram of an image processing module according to at least one exemplary embodiment is shown.

[0012] Figure 2B An encoder system according to at least one exemplary embodiment is illustrated.

[0013] Figure 3A The figure illustrates a decoder system according to at least one exemplary embodiment.

[0014] Figure 3B A block diagram of a projector module according to at least one exemplary embodiment is shown.

[0015] Figure 4 The illustration shows a block diagram of a method for conducting a web-based video call, according to at least one exemplary embodiment.

[0016] Figure 5 The diagram illustrates a block diagram of another part of a method for conducting a web-based video call according to at least one exemplary embodiment.

[0017] Figure 6 Examples of computer devices and mobile computer devices according to at least one exemplary embodiment are illustrated.

[0018] It should be noted that these figures are intended to illustrate the general characteristics of the methods, structures, and / or materials used in some exemplary embodiments and to supplement the written description provided below. However, these figures are not drawn to scale and may not accurately reflect the precise structural or performance characteristics of any given embodiment, and should not be construed as limiting or restricting the range of values ​​or properties covered by the exemplary embodiments. For example, the relative thickness and positioning of molecules, layers, regions, and / or structural elements may be reduced or enlarged for clarity. The use of similar or identical reference numerals in the various figures is intended to indicate the presence of similar or identical elements or features. Detailed Implementation

[0019] The user experience in video calls can be less than ideal because participants are presented in their own environments and are restricted in their display (e.g., rectangular display). Participants in different environments can cause participants to feel socially distant and / or lead to unwanted social interactions between users.

[0020] To address the aforementioned issues, an image of a first participant in a video call can be extracted from a first environment and projected (or reprojected) into (or re-projected into) a second participant's second environment. In other words, an exemplary embodiment can generate images of one or more participants viewed on a device (e.g., a mobile device) and transmit them to another environment (e.g., the environment of another participant). The implementation enables virtual teleportation video calls, including generating and transmitting an image of a first video call participant, projecting the image of the first video call participant into the field of view of a second video call participant's device, and allowing the second video call participant to walk towards and / or move around the first video call participant as if the first video call participant were in the second video call participant's space.

[0021] Exemplary implementations can include segmenting the streaming video of at least one participant and projecting the segmented portion, including the participant's streaming video, onto the device's display. Furthermore, the generation and transmission of the participant's image can be implemented in real-time (e.g., live, with minimal latency) within a webpage. The webpage implementation may not involve installing the application on a local device. In other words, the implementation can operate in a zero-installation computing environment (e.g., without users downloading files or inserting storage to install the application). Zero-installation computing environments offer the advantage of flexibility because changes to a web-based video calling application affect all users of the web-based video calling application, and users do not need to take any action other than opening a webpage to use it.

[0022] Additionally, video calling applications can be web-based and / or use applications installed on the device. In either case, a server configured to control streaming communication is used. In an exemplary embodiment, video calling can stream video directly from a first device to a second device. Therefore, exemplary embodiments disclose novel video calling features that at least include providing tools for virtually remotely transmitting video calls in a browser without using a server (e.g., a third-party server) configured to control streaming communication. In an exemplary embodiment, web-based refers to the functionality implemented in a browser using, for example, a web server to transmit images, videos, text, etc., that can be displayed as web pages in the browser on the display of a computing device (using HTTP via the Internet). Furthermore, the web server can transmit software code (e.g., JavaScript, C++, Visual Basic, etc.) that can be executed by the computing device in association with the web page. The server configured to control streaming communication (e.g., a third-party server) is (or operationally) independent of (e.g., different from) a web server (or operates independently of a web server).

[0023] Figure 1 A block diagram illustrating a signal flow according to at least one exemplary embodiment is shown. Figure 1 As shown, signal stream 100 includes a capture block 110, a segment block 115, a communication block 125, a projector block 130, and a rendering block 135. In the capture block 110, images 105 are captured using a camera of a computing device (e.g., a desktop computer, laptop computer, mobile device, stand-alone image capture system, etc.). Through the embodiments shown herein, images of one or more participants can be projected (or reprojected) from at least one environment to at least one other environment viewed on the device. Therefore, signal stream 100 can illustrate an exemplary implementation of virtual teleportation video calls.

[0024] Image 105 may be a frame of video corresponding to a video call on a first device 145 including a camera. Image 105 may include pixels corresponding to at least the participants in the video call and the environment (sometimes referred to as the background) in which the participants are located. Data representing image 105 is transmitted to segment 115 blocks. Segment 115 blocks may divide image 105 into at least two segments, one of which may be pixels corresponding to the participants in the video call. The segmented image including the participants in the video call is shown as image 120.

[0025] Image 120 is transmitted from a first device 145 (e.g., a device including capture blocks 110 and 115) to a second device 150 via communication block 125. Communication block 125 is capable of streaming video corresponding to a video call, where image 120 represents a frame of video. Image 120 can be a portion of a captured frame (e.g., image 105). Therefore, image 120 can include less data than a complete frame. Thus, the exemplary implementation can use fewer resources (e.g., bandwidth) than streaming a complete frame of video. Streaming can use web-based communication standards (e.g., WebRTC, VoIP, RTP, PTP telephony, etc.).

[0026] The second device is capable of receiving image 120 and projecting (projector 130) the image onto an image captured by the second device that is not yet displayed on the second device (e.g., a real-world image). Projecting image 120 can include selecting the position and orientation of image 120 relative to the image captured by the second device. The resulting image can be rendered (rendering block 135) and displayed on the second device, as shown in image 140.

[0027] Figure 2A A block diagram of an image processing module according to at least one exemplary embodiment is illustrated. Figure 2A As shown, the image processing module 230 includes an object recognizer 235, a segmenter 240, an image modification module 245, and a fragment 115. As described above, video calls can be implemented within a webpage. Therefore, the image processing module 230 can be an element of a webpage. For example, the image processing module 230 can be implemented using JavaScript. The image processing module 230 can include machine learning elements; for example, it can include a machine learning model (e.g., a convolutional neural network (CNN)). Therefore, the image processing module 230 can include a trained machine learning (ML) model implemented in JavaScript (e.g., TensorFlow.js). Furthermore, the image processing module 230 can be loaded onto a computing device along with the webpage. Therefore, the trained ML model implemented in JavaScript can be loaded onto the computing device along with the webpage. Therefore, the image processing module 230 can enable (or help enable) virtual remote video calls in a browser without installing the application on the device.

[0028] Object recognizer 235 can be configured to recognize each object in an image or frame of a video call. Object recognizer 235 can be configured to identify an object as a participant in a video call. Object recognizer 235 can use a trained ML model (e.g., a convolutional neural network (CNN)) to recognize objects; therefore, object recognizer 235 can use a trained ML model implemented in JavaScript (e.g., TensorFlow.js) to recognize objects.

[0029] A video image or frame (e.g., image 105) can include multiple objects. A trained ML model associated with object identifier 235 can place multiple boxes (sometimes referred to as bounding boxes) on the image. The object identifier can associate data (e.g., features associated with pixels of the image) with each box. The data can indicate the object in the box (the object can be a blank or part of an object). The object can be identified by its features. The accumulated data is sometimes referred to as a class or classifier. The class or classifier can be associated with the object. The data (e.g., bounding boxes) can also include confidence scores (e.g., numbers between zero (0) and one (1)).

[0030] After a trained ML model processes images or frames of a video, it can handle multiple classifiers that indicate an object or a part of an object. In other words, an object (or a part of an object) can be within multiple overlapping bounding boxes. However, the confidence scores of each classifier can be different. For example, a classifier that identifies a part of an object can have a lower confidence score than a classifier that identifies the whole (or substantially whole) object. The trained ML model can also be configured to discard bounding boxes that do not have an associated classifier. In other words, the trained ML model can discard bounding boxes that do not contain objects. The trained ML model can then use the classifier with the highest confidence score to identify an object. One of the objects can be identified as a participant in a video call (e.g., as a human or part of a human).

[0031] Segmenter 240 can be configured to generate an image that includes the participants (and has no other objects or background pixels). For example, object identifier 235 can pass the coordinates of the bounding boxes that include the participants. Segmenter 240 can remove pixels in the image that are not within the bounding boxes. Segmenter 240 can copy the contents of the bounding boxes to a new image. Furthermore, segmenter 240 can be configured to modify the boundaries of objects to remove any unwanted pixels and smooth transitions to improve the image as a participant (e.g., image 120). The segmented image is stored as fragment 115.

[0032] In some implementations, the object recognizer 235 and the segmenter 240 can be combined into a single operation. For example, image segmentation for body parts can be part of an ML tool or model. This ML tool can be configured to group pixels in an image into semantic regions to locate objects and boundaries. For example, the ML tool or model can be configured to classify pixels in an image into two categories: 1) pixels representing people and 2) pixels representing background. Then, pixels representing people can be segmented from the image.

[0033] The image modification module 245 is configured to modify the segmented image and / or generate a new image (e.g., as segment 115) based on the segmented image. The image modification module 245 can be configured to generate a three-dimensional (3D) image from a two-dimensional image. The 2D-3D conversion tool can be an element of a webpage (e.g., as JavaScript). The conversion tool can be implemented using the segmented image as input. For example, the conversion tool can use (e.g., a depth map associated with the participant), a 3D mesh, a deformation algorithm, etc. The conversion tool can be implemented as a trained ML model. In an exemplary embodiment, the 2D-to-3D conversion can be a partial conversion (e.g., adding depth to a portion of the segmented image).

[0034] The image modification module 245 can be configured to apply image filters to the segmented image. For example, the image modification module 245 can apply ghosting filters, color filters, enhancement filters, holographic filters, overlay filters, etc. The image modification module 245 can be configured to enhance the segmented image (e.g., improve quality, resolution, etc.). The image modification module 245 can be configured to complete the segmentation of the image. For example, the segmented image can be a part of a participant (e.g., a head), and the image modification module 245 can add to the segmented image (e.g., add a body). For simplicity, the image modification module 245 can be configured to modify the segmented image using other techniques not described herein.

[0035] exist Figure 2B In the examples, encoder system 200 may be or include at least one computing device, and should be understood to represent virtually any computing device configured to perform the techniques described herein. Therefore, encoder system 200 can be understood to include various components that can be used to implement the techniques described herein or different or future versions thereof. As an example, encoder system 200 is shown to include at least one processor 205 and at least one memory 210 (e.g., a non-transitory computer-readable storage medium).

[0036] Figure 2B An encoder system according to at least one exemplary embodiment is illustrated. Figure 2BAs shown, the encoder system 200 includes at least one processor 205, at least one memory 210, a controller 220, and an encoder 225. The at least one processor 205, at least one memory 210, controller 220, and encoder 225 are communicatively coupled via bus 215. The encoder system can be an element of a video call implemented via a webpage. In an exemplary embodiment, the encoder 225 and controller 220 are loaded onto a computer when a webpage configured to implement a video call is loaded. The encoder 225 and controller 220 can use web-based communication standards (e.g., WebRTC, VoIP, RTP, PTP telephony, etc.) (or elements thereof). The encoder system 200 can use segment 115 as input.

[0037] At least one processor 205 can be used to execute instructions stored on at least one memory 210. Therefore, at least one processor 205 is capable of implementing the various features and functions described herein, or additional or alternative features and functions. For example, processor 205 is capable of executing code associated with a webpage configured to implement a video call stored in at least one memory 210. At least one processor 205 and at least one memory 210 can be used for a variety of other purposes. For example, at least one memory 210 can represent various types of memory that can be used to implement any of the modules described herein, as well as examples of associated hardware and software.

[0038] At least one memory 210 may be configured to store data and / or information associated with the encoder system 200 (e.g., to implement web-based communication standards (e.g., WebRTC, VoIP, RTP, PTP telephony, etc.)). At least one memory 210 may be a shared resource. For example, the encoder system 200 may be a component of a larger system (e.g., a server, personal computer, mobile device, etc.). Therefore, at least one memory 210 may be configured to store data and / or information associated with other components within the larger system (e.g., image / video services, web browsing, or wired / wireless communication).

[0039] Controller 220 can be configured to generate various control signals and transmit them to various blocks in encoder system 200 and / or image processing 230 modules. Controller 220 can be configured to generate control signals to implement the techniques described herein. According to an exemplary embodiment, controller 220 can be configured to control encoder 225 to encode images, image sequences, video frames, video frame sequences, etc. For example, controller 220 can generate control signals corresponding to the encoding and transmission of images (or frames) associated with web-based communication standards (e.g., WebRTC, VoIP, RTP, PTP telephony, etc.).

[0040] Encoder 225 can be configured to receive input image 5 (and / or video stream) and output compressed (e.g., encoded) bits 10. Encoder 225 can convert the video input into discrete video frames (e.g., as images). Input image 5 can be compressed (e.g., encoded) into compressed image bits. Encoder 225 can further convert each image (or discrete video frame) into a matrix of blocks or macroblocks (hereinafter referred to as blocks). For example, images can be converted into blocks of 32×32, 32×16, 16×16, 16×8, 8×8, 4×8, 4×4, or 2×2 matrices, each block having multiple pixels. Although eight (8) exemplary matrices are listed, exemplary implementations are not limited thereto.

[0041] The compressed bit 10 may represent the output of the encoder system 200. For example, the compressed bit 10 may represent an encoded image (or video frame). For example, the compressed bit 10 may be stored in a memory (e.g., at least one memory 210). For example, the compressed bit 10 may be ready for transmission to a receiving device (not shown). For example, the compressed bit 10 may be sent to a system transceiver (not shown) for transmission to the receiving device.

[0042] At least one processor 205 may be configured to execute computer instructions associated with the image processing 230 module, the controller 220, and / or the encoder 225. At least one processor 205 may be a shared resource. For example, the encoder system 200 may be a component of a larger system (e.g., a mobile device, desktop computer, laptop computer, etc.). Therefore, at least one processor 205 may be configured to execute computer instructions associated with other components within the larger system (e.g., image / video capture, web browsing, and / or wired / wireless communication).

[0043] In an exemplary embodiment, the image processing module 230 can be an element of the encoder 225. For example, image 5 can be multiple frames of a streaming video. The encoder 225 can be configured to process each frame individually. Thus, the encoder can select frames of the streaming video and transmit the selected frames to the image processing module 230 as input to the image processing module. After processing the frames, the image processing module 230 can generate segment 115, which can then be processed (e.g., compressed) by the encoder 225.

[0044] Figure 3A A decoder system according to at least one exemplary embodiment is illustrated. Figure 3A As shown, the decoder system 300 includes at least one processor 305, at least one memory 310, a controller 320, and a decoder 325. The at least one processor 305, at least one memory 310, controller 320, and decoder 325 are communicatively coupled via bus 315.

[0045] exist Figure 3A In the examples, decoder system 300 may be at least one computing device and should be understood to represent, in effect, any computing device configured to perform the techniques described herein. Therefore, decoder system 300 can be understood to include various components that can be used to implement the techniques described herein or different or future versions thereof. For example, decoder system 300 is shown as including at least one processor 305 and at least one memory 310 (e.g., a computer-readable storage medium).

[0046] Therefore, at least one processor 305 can be used to execute instructions stored on at least one memory 310. Thus, at least one processor 305 is capable of implementing the various features and functions described herein, or additional or alternative features and functions (e.g., to implement web-based communication standards (e.g., WebRTC, VoIP, RTP, PTP telephony, etc.)). At least one processor 305 and at least one memory 310 can be used for a variety of other purposes. For example, at least one memory 310 can be understood to represent examples of various types of memory and associated hardware and software that can be used to implement any of the modules described herein. According to exemplary embodiments, the decoder system 300 can be included in a larger system (e.g., a personal computer, a laptop computer, a mobile device, etc.).

[0047] At least one memory 310 may be configured to store data and / or information associated with the projector 130 and / or the decoder system 300. At least one memory 310 may be a shared resource. For example, the decoder system 300 may be a component of a larger system (e.g., a personal computer, mobile device, etc.). Therefore, at least one memory 310 may be configured to store data and / or information associated with other components within the larger system (e.g., web browsing or wireless communication) (e.g., to implement web-based communication standards (e.g., WebRTC, VoIP, RTP, PTP telephony, etc.)).

[0048] The controller 320 can be configured to generate various control signals and transmit these control signals to various blocks in the projector and / or decoder system 300. The controller 320 can be configured to generate control signals to implement the video encoding / decoding techniques described herein. According to an exemplary embodiment, the controller 320 can be configured to control the decoder 325 to decode video frames.

[0049] Decoder 325 can be configured to receive compressed (e.g., encoded) bits 10 as input and output image 5. The compressed (e.g., encoded) bits 10 can also represent compressed video bits (e.g., video frames). Therefore, decoder 325 can convert discrete video frames of compressed bits 10 into a video stream.

[0050] At least one processor 305 may be configured to execute computer instructions associated with the projector 130, controller 320, and / or decoder 325. At least one processor 305 may be a shared resource. For example, the decoder system 300 may be a component of a larger system (e.g., a personal computer, mobile device, etc.). Therefore, at least one processor 305 may be configured to execute computer instructions associated with other components within the larger system (e.g., web browsing or wireless communication).

[0051] Figure 3B A block diagram of a projector 130 according to at least one exemplary embodiment is illustrated. Figure 3B As shown, the projector 130 includes a plane generator 330 module, a plane locator 335 module, a normal determination 340 module, and a projection 345 module. In an exemplary embodiment, a web-based video call includes at least two computing devices. The first computing device can be used by a first participant, while the second computing device can be used by a second participant. In a web-based virtual teleportation video call, the first device can include a reference... Figure 2A and Figure 2B The described elements are used to generate an image of the first participant, while the second device is capable of including references. Figure 3A and Figure 3B The described elements are configured to receive an image of a first participant and project the first participant into the environment of a second participant. Therefore, the projector 130 advantageously enables (or helps enable) virtual teleportation video calls to be performed in a browser without installing an application on the device. Furthermore, exemplary embodiments can include transmitting images of the first participant, the second participant, and / or both the first and second participants.

[0052] Thus, the plane generator 330 module can be configured to generate a plane to project the image of the first participant onto the second device. The plane locator 335 module can be configured to select a position on the display of the second device to place the plane in the physical world environment of the second device. In an exemplary embodiment, the user of the second device can reference the real-world position to be rendered on the display of the second device. This is sometimes referred to as mixed reality. The normal determination 340 module can be configured to orient the plane on the display of the second device. The projection 345 module can be configured to project the first participant (e.g., as fragment 115) onto the plane.

[0053] The plane generator 330 module can generate a plane by generating a 2D virtual structure (e.g., a rectangle). The 2D structure can have a size based on its proximity to the plane in the real world. In some cases, the rectangle may be larger than the device itself if the user is close, and only a sub-part of the first participant may be displayed. The 2D structure can have a size based on the display of a second device. For example, the 2D structure can have a size proportional to the display and smaller than the display. The 2D structure can be transparent (e.g., to allow the background to be visible). The 2D structure can be implemented via function calls (e.g., via web-based display tools and / or web browsers). An exemplary code snippet is shown below:

[0054] marker.setAttribute('position',{

[0055] x:cursor.intersection.point.x,

[0056] y:cursor.intersection.point.y+0.5,

[0057] z:cursor.intersection.point.z});

[0058] var rot=cam.getAttribute('rotation');

[0059] marker.setAttribute('rotation',{

[0060] x:0,

[0061] y:rot.y,

[0062] z:0

[0063] });

[0064] The plane locator 335 module can be configured to select a position on the display of a second device to place the plane. This position can be based on an image (e.g., a preview image) captured by the second device and displayed on the device using a web-based application (e.g., a browser). For example, see reference... Figure 1Image 140 is an image of a corridor. The plane's position can be placed approximately centrally within the corridor and at a comfortable viewing depth. Therefore, the position can be based on pixel position (e.g., X, Y position) and depth. Depth can be determined using a depth sensor or camera on the device and / or calculated using a depth algorithm (e.g., using a web-based tool, a web-based augmented reality tool, and / or a JavaScript tool (e.g., WebXR)). In other words, depth can be determined using function calls in a web-based (e.g., JavaScript) augmented reality tool (e.g., WebXR) capable of returning a depth map. In an exemplary implementation, a web-based tool can be used to determine the position, which is configured to render virtual objects (e.g., a plane or communication image of a participant) in the real world (e.g., as an image captured using the device's camera and displayed via a web browser). An exemplary code snippet is shown below:

[0065] var sc=document.querySelector('a-scene');

[0066] var cam=document.getElementsByTagName('a-camera')[0];

[0067] var cursor=sc.querySelector('[ar-raycaster]').components.cursor; if(cursor.intersection){

[0068] }

[0069] The normal determination module 340 can be configured to orient a plane on the display of a second device. For example, it can determine the normals associated with the real world (e.g., as an image captured using the device's camera and displayed via a web browser) and the plane's position. The plane can then be oriented (e.g., translated, rotated, etc.) relative to the real world such that the plane is approximately perpendicular to the normals. In an exemplary implementation, the normal orientation associated with each pixel in the real world can be determined (e.g., estimated). Once a normal vector is associated with each pixel, the normal can be associated with the pixel in world coordinates. Web-based augmented reality tools and / or JavaScript tools (e.g., WebXR) can be used to generate the normal vectors and pixels. For example, the normal vectors can be projected from the plane onto pixels in the real world. Orientation can be achieved and / or confirmed by projecting the normal vectors from the orienting plane into the real world. The normal vectors should be approximately equal and in opposite directions. An example code snippet is shown below:

[0070] var rot=cam.getAttribute('rotation');

[0071] marker.setAttribute('rotation',{

[0072] x:0,

[0073] y:rot.y,

[0074] z:0

[0075] });

[0076] The projection module 345 can be configured to project a first participant (e.g., as fragment 115) onto a plane. For example, pixels representing an image of the first participant (e.g., decompressed fragment 115) can be added to the plane. The projection can include adding appearance and feel features to improve the user experience. For example, shadows can be added to the resulting image. The projection module 345 can be implemented to render (e.g., using an image transmitted by the first participant) a modified real-world image on the display of a second device. Rendering can be a function of displaying a webpage in a browser that implements video calling and / or web-based communication standards (e.g., WebRTC, VoIP, RTP, PTP telephony, etc.). An exemplary code snippet is shown below:

[0077] marker.setAttribute('src',data.src);

[0078] Figure 4 and Figure 5 The diagram illustrates a block diagram of the method. This can be achieved by executing data stored in a device (e.g., such as...). Figure 2B and 3A The reference is executed by software code in the associated memory (e.g., at least one memory 210, 310) and by at least one processor (e.g., at least one processor 205, 305) associated with the device. Figure 4 and Figure 5 The steps described herein. However, alternative embodiments are contemplated, such as systems embodied as dedicated processors. Although the steps described below are described as being performed by a processor, these steps are not necessarily performed by the same processor. In other words, at least one processor can perform the steps described in the following references. Figure 4 and Figure 5 The steps described.

[0079] Figure 4 The illustration shows a block diagram of a method for conducting a web-based video call, according to at least one exemplary embodiment. Figure 4 As shown, in step S405, a web-based video conference is established. A web-based video call can be a virtual remote video call implemented in a browser. For example, a first participant on a first device can open a webpage containing a video calling web application and use the video calling web application (or other communication tools (e.g., email, messaging, etc.)) to invite a second participant on a second device. The second participant can join the video call by opening a webpage containing the video calling web application on the second device and requesting to join. The first participant can accept the second participant's entry into the video call. For example, a first participant on a first device can open a webpage containing a video calling web application and use the video calling web application to call a second participant on a second device. The second participant can join the video call by opening a webpage containing the video calling web application to answer the video call.

[0080] Video calling applications can be web-based and / or use applications installed on the device. In either case, a server configured to control streaming communication is used. In an exemplary embodiment, video calling can stream video directly from a first device to a second device. In other words, video calling can be streamed from a first device to a second device (e.g., end-to-end) without using a server configured to control streaming communication (e.g., a third-party server). Therefore, exemplary embodiments implement new video calling functionality (e.g., virtual telepresence video calls) in a browser without using a server configured to control streaming communication (e.g., a third-party server).

[0081] By enabling video calls in a browser, exemplary implementations enable (or help enable) zero-installation (e.g., no user-installed applications or plugins) video calls. By eliminating the server, exemplary implementations enable (or help enable) communications with low latency, device / platform independence (e.g., working in any browser), enhanced security (e.g., the server can add a security risk layer without adding third-party services), and adaptability to network conditions, without requiring dedicated tools (e.g., plugins). Web pages including video calling web applications can use web-based communication standards (e.g., WebRTC, VoIP, RTP, PTP telephony, etc.).

[0082] In step S410, video is captured. For example, the first participant can use a computing device (e.g., a desktop computer, laptop computer, mobile device, etc.) to capture video. The video can consist of multiple frames. Each frame can be used in a real-time virtual remote video call via a webpage executed in a browser. Each frame can represent the first participant as an image of the person to be transmitted.

[0083] In step S415, the video is segmented. For example, each frame of a streaming video can be segmented. Image segmentation for body parts can be part of an ML tool or model. This ML tool can be configured to group pixels in an image into semantic regions to locate objects and boundaries. For example, the ML tool or model can be configured to classify pixels in an image into two categories: 1) pixels representing people and 2) pixels representing background. Then, pixels representing people can be segmented from the image. Alternatively, the segmented frame can include identifying each object in the frame and selecting the first participant as a segment. For example, as discussed in more detail above, a trained ML model can place multiple boxes (sometimes called bounding boxes) on the image. An object identifier can associate data (e.g., features associated with pixels in the image) with each box. The data can indicate the object in the box (the object can be empty or part of an object). The object can be identified by its features. The features can be classified as people. Then, pixels representing people can be segmented from the image.

[0084] Web pages (e.g., video calling web applications running in a browser) can include trained machine learning (ML) models implemented in JavaScript (e.g., Tensorflow.js). These trained ML models implemented in JavaScript can be configured to segment frames (e.g., identify objects) and select the first participant (e.g., a human) as the image segment (e.g., segment 115).

[0085] In an exemplary implementation, there may be two or more participants (e.g., humans). The first participant may not be the one standing closer to the camera of the first device and in the full view that is to be projected. Therefore, one to n participants can be selected from the scene as needed to project one to n people simultaneously.

[0086] In step S420, at least one segment is processed. In an exemplary embodiment, processing at least one segment can be optional. In other words, processing can continue to step S425 without performing step S420. Processing at least one segment can include image modification of at least one segment (e.g., segment 115). For example, image modification can include image enhancement (e.g., quality improvement), image conversion (e.g., 2D to 3D), image warping (making a 2D image appear 3D without 3D conversion), etc.

[0087] In step S425, at least one segment is encoded. For example, at least one segment can be encoded using a standard for video calling. The standard for video calling can include loading an encoder when establishing a web-based video conference. For example, WebRTC can be used for video calling. WebRTC-compatible browsers may or may not use or support at least the VP8 and / or AVC encoder / decoder standards. In step S430, at least one encoded segment is streamed. For example, at least one encoded segment can be transmitted from a first device to a second device via, for example, the Internet using the WebRTC standard. In another embodiment, data (e.g., raw binary data) can be transmitted without using a standard.

[0088] Figure 5 The diagram illustrates a block diagram of another part of a method for conducting a web-based video call, according to at least one exemplary embodiment. Figure 5 As shown, in step S505, a web-based video conference is established. For example, a first participant on the first device can open a webpage containing a video calling web application and invite a second participant on the second device using the video calling web application (or other communication tools (e.g., email, messaging, etc.)). The second participant can join the video call by opening a webpage containing the video calling web application on the second device and requesting to join. The first participant can accept the second participant's entry into the video call. For example, the first participant on the first device can open a webpage containing a video calling web application and call a second participant on the second device using the video calling web application. The second participant can join the video call by opening a webpage containing the video calling web application to answer the video call. The video call can stream video directly from the first device to the second device. In other words, the video call can be streamed from the first device to the second device without using a server configured to control the streaming communication.

[0089] Video calling applications can be web-based and / or use applications installed on the device. In either case, a server configured to control streaming communication is used. In an exemplary embodiment, video calling can stream video directly from a first device to a second device. In other words, video calling can be streamed from a first device to a second device (e.g., end-to-end) without using a server configured to control streaming communication (e.g., a third-party server). Therefore, exemplary embodiments implement new video calling functionality (e.g., virtual telepresence video calls) in a browser without using a server configured to control streaming communication (e.g., a third-party server).

[0090] By enabling video calls in a browser, exemplary implementations enable (or help enable) zero-installation (e.g., no user-installed applications or plugins) video calls. By eliminating the server, exemplary implementations enable (or help enable) communications with low latency, device / platform independence (e.g., working in any browser), enhanced security (e.g., the server can add a security risk layer without adding third-party services), and adaptability to network conditions, without requiring dedicated tools (e.g., plugins). Web pages including video calling web applications can use web-based communication standards (e.g., WebRTC, VoIP, RTP, PTP telephony, etc.).

[0091] In step S510, a video stream is received. For example, encoded frames of video corresponding to a video call can be transmitted from a first device to a second device via, for example, the Internet using the WebRTC standard. The second device can receive the video stream frame by frame and / or in groups of frames. In step S515, the video stream is decoded into a first video. For example, each frame of the streaming video can be decoded. The standard used for making video calls can include loading a decoder when establishing a web-based video conference. For example, WebRTC can be used for making video calls. WebRTC-compatible browsers can use or support at least the VP8 and / or AVC encoder / decoder standards. In another embodiment, data (e.g., raw binary data) can be transmitted without using a standard.

[0092] In step S520, video is captured as a second video. For example, the second participant can use a computing device (e.g., a desktop computer, laptop computer, mobile device, etc.) to capture the video. The video can consist of multiple frames. Each frame can be used in a real-time virtual remote video call via a webpage executed in a browser. Each frame can represent a real-world image of the first participant that can be projected onto it.

[0093] In step S525, a normal vector associated with the second video is determined. For example, a normal vector associated with the real world (e.g., an image captured using the device's camera and displayed via a web browser) and a planar position can be determined. In an exemplary implementation, a normal vector associated with each pixel in the real world can be determined (e.g., estimated). Once a normal vector is associated with each pixel, it can be associated with the pixel in world coordinates. Web-based augmented reality tools and / or JavaScript tools (e.g., WebXR) can be used to generate the normal vector and pixels. For example, the normal vector can be projected from a plane onto pixels in the real world.

[0094] In step S530, a plane is generated. For example, the plane can be generated by generating a 2D virtual structure (e.g., a rectangle). The 2D structure can have a size based on the display of the second device. For example, the 2D structure can have a size that is proportional to the display and smaller than the display. The 2D structure can be transparent (e.g., to allow the background to be visible). The 2D structure can be implemented via function calls (e.g., via a web-based display tool and / or a web browser).

[0095] In step S535, the plane is oriented based on the normal vector. For example, the plane can be oriented relative to the real world (e.g., translated, rotated, etc.) such that the plane is approximately oriented perpendicular to the normal vector. Orientation can be achieved and / or confirmed by projecting the normal vector from the plane into the real world. The normal vector associated with the real world and the normal vector associated with the plane should be approximately equal and in opposite directions.

[0096] In step S540, the first video is projected onto the plane of the second video. For example, pixels representing the image of the first participant (e.g., decompressed segment 115) can be added to the plane. In step S545, the first and second videos are rendered. The projection can be implemented as rendering a modified second (or real-world) video (e.g., using the first video (or the transmitted image of the first participant)) on the display of a second device. Rendering can be the function of displaying a webpage in a browser that implements video calling and / or web-based communication standards (e.g., WebRTC, VoIP, RTP, PTP telephony, etc.).

[0097] Figure 6 Examples of computer devices 600 and 650 that can be used with the techniques described herein are shown. Computer device 600 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computer device 650 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the embodiments of the invention described and / or claimed in this document.

[0098] Computing device 600 includes a processor 602, a memory 604, a storage device 606, a high-speed interface 608 connected to the memory 604 and a high-speed expansion port 610, and a low-speed interface 612 connected to a low-speed bus 614 and the storage device 606. Each of components 602, 604, 606, 608, 610, and 612 is interconnected using various buses and may be suitably mounted on a common motherboard or otherwise mounted. Processor 602 is capable of processing instructions for execution within computing device 600, including instructions stored in memory 604 or storage device 606, to display graphical information of a GUI on an external input / output device, such as a display 616 coupled to high-speed interface 608. In other embodiments, multiple processors and / or multiple buses, as well as multiple memories and various types of memory, may be suitably used. Furthermore, multiple computing devices 600 may be connected, with each device providing a portion of the necessary operation (e.g., as a server group, a set of blade servers, or a multiprocessor system).

[0099] Memory 604 stores information within computing device 600. In one embodiment, memory 604 is one or more volatile memory cells. In another embodiment, memory 604 is one or more non-volatile memory cells. Memory 604 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.

[0100] Storage device 606 provides large-capacity storage for computing device 600. In one embodiment, storage device 606 may be or include computer-readable media, such as floppy disk devices, hard disk devices, optical disk devices, magnetic tape devices, flash memory or other similar solid-state storage devices or arrays of devices, including devices in storage area networks or other configurations. A computer program product can be tangibly embodied in an information carrier. The computer program product may also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as memory 604, storage device 606, or memory on processor 602.

[0101] High-speed controller 608 manages bandwidth-intensive operations of computing device 600, while low-speed controller 612 manages less bandwidth-intensive operations. This functional allocation is merely exemplary. In one embodiment, high-speed controller 608 is coupled to memory 604, display 616 (e.g., via a graphics processor or accelerator), and high-speed expansion port 610, which can accept various expansion cards (not shown). In this embodiment, low-speed controller 612 is coupled to storage device 606 and low-speed expansion port 614. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, Wireless Ethernet), may be coupled to one or more input / output devices, such as keyboards, pointing devices, scanners, or networking devices such as switches or routers, for example, via a network adapter.

[0102] Computing device 600 can be implemented in a variety of different forms, as shown in the figure. For example, it can be implemented as a standard server 620, or multiple times in a set of such servers. It can also be implemented as part of a rack server system 624. Alternatively, it can be implemented in a personal computer such as a laptop computer 622. Alternatively, components from computing device 600 can be combined with other components in mobile devices (not shown), such as device 650. Each of these devices can contain one or more of computing devices 600, 650, and the entire system can consist of multiple computing devices 600, 650 communicating with each other.

[0103] Computing device 650 includes processor 652, memory 664, input / output devices such as display 654, communication interface 666 and transceiver 668, and other components. Device 650 may also be equipped with storage devices, such as microdrives or other devices, to provide additional storage. Each of components 650, 652, 664, 654, 666, and 668 is interconnected using various buses, and some components may be mounted on a common motherboard or otherwise suitably mounted.

[0104] Processor 652 is capable of executing instructions within computing device 650, including instructions stored in memory 664. The processor can be implemented as a chipset comprising individual and multiple analog and digital processors. The processor can provide, for example, coordination of other components of device 650, such as control of the user interface, applications running by device 650, and wireless communications performed by device 650.

[0105] Processor 652 can communicate with the user via control interface 658 and display interface 656 coupled to display 654. Display 654 can be, for example, a TFT LCD (Thin Film Transistor Liquid Crystal Display) or OLED (Organic Light Emitting Diode) display or other suitable display technology. Display interface 656 can include suitable circuitry for driving display 654 to present graphics and other information to the user. Control interface 658 can receive commands from the user and translate them for submission to processor 652. Additionally, an external interface 662 can be provided to communicate with processor 652 to enable near-field communication between device 650 and other devices. External interface 662 can provide wired communication in some embodiments, wireless communication in others, and multiple interfaces can also be used.

[0106] Memory 664 stores information within computing device 650. Memory 664 can be implemented as one or more computer-readable media, one or more volatile memory cells, or one or more non-volatile memory cells. An expansion memory 674 may also be provided and connected to device 650 via an expansion interface 672, which may include, for example, a SIMM (Single In-line Memory Module) card interface. Such an expansion memory 674 can provide additional storage space for device 650, or it may also store applications or other information of device 650. Specifically, expansion memory 674 may include instructions for performing or supplementing the above processes, and may also include security information. Thus, for example, expansion memory 674 may be provided as a security module of device 650 and can be programmed with instructions that allow secure use of device 650. Furthermore, secure applications and additional information, such as placing identification information on the SIMM card in a hackable manner, may be provided via a SIMM card.

[0107] The memory may include, for example, flash memory and / or NVRAM memory, as described below. In one embodiment, the computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium that can be received, for example, via transceiver 668 or external interface 662, such as memory 664, extended memory 674, or memory on processor 652.

[0108] Device 650 can communicate wirelessly via communication interface 666, which may include digital signal processing circuitry if necessary. Communication interface 666 can provide communication under various modes or protocols, such as GSM voice calls, SMS, EMS or MMS messages, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS. This communication can occur, for example, via radio frequency transceiver 668. Furthermore, short-range communication can occur, such as using Bluetooth, Wi-Fi, or other transceivers (not shown). Additionally, GPS (Global Positioning System) receiver module 670 can provide device 650 with additional navigation and location-related wireless data, which can be appropriately used by applications running on device 650.

[0109] Device 650 can also communicate audibly using audio codec 660, which can receive spoken information from a user and convert it into usable digital information. Audio codec 660 can also generate audible sound for the user, such as through a speaker (e.g., in the handset of device 650). Such sound can include sounds from voice telephone calls, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by applications operating on device 650.

[0110] The computing device 650 can be implemented in a variety of different forms, as shown in the figure. For example, it can be implemented as a cellular phone 680. It can also be implemented as part of a smartphone 682, a personal digital assistant, or other similar mobile device.

[0111] In general, an apparatus, system, non-transitory computer-readable medium (on which computer-executable program code capable of being executed on a computer system) and / or method can perform a process by means of: opening a web-based video call in a browser on a first device; receiving a request to join the web-based video call from a second device; capturing video including frames by the first device; segmenting the frames by the first device; selecting at least one segment of the segmented frames by the first device; and streaming the video including the at least one segment directly from the first device to the second device as a real-time virtual teletransmission.

[0112] The implementation may include one or more of the following features. For example, opening the web-based video call includes loading a webpage containing code configured to implement a trained machine learning model, which can be configured to segment the frame and select the at least one segment. The at least one segment may be an image of a participant in the web-based video call. The web-based video call may be implemented using a web-based communication standard. The frame segmentation may include: grouping pixels in the frame into semantic regions to locate objects and boundaries, classifying the pixels of the frame into two categories: 1) pixels representing people and 2) pixels representing background, and segmenting the pixels representing people from the frame. The frame segmentation may include identifying each object in the frame, selecting the at least one segment includes selecting an object as the at least one segment, and the object may be a participant in the web-based video call. The at least one segment may be an image of a participant in the web-based video call, and the method may further include converting the image from a two-dimensional image to a three-dimensional image. The at least one segment may be an image of a participant in the web-based video call, and the method may further include applying a filter to the image. The web-based video call can be implemented with zero web application installation.

[0113] In another general aspect, an apparatus, system, non-transitory computer-readable medium (on which computer-executable program code capable of being executed on a computer system) and / or method can perform a process by means of: opening a web-based video call webpage in a browser on a first device; transmitting a request from the first device to join a web-based video call from a second device; receiving streaming video directly from the second device at the first device as a first video; capturing a second video by the first device; orienting the first video based on the second video by the first device; projecting the first video onto the second video by the first device to generate a third video; and rendering a webpage including the third video by the first device.

[0114] The implementation may include one or more of the following features. For example, the method may further include generating a plane, and the orientation of the first video may include: determining a normal vector associated with the second video, and rotating and translating at least one of the first videos based on the normal vector. The method may further include generating a plane and positioning the plane in the second video, wherein projecting the first video into the second video includes adding the first video to the plane. The method may further include generating a plane and positioning the plane in the second video, and the orientation of the first video may include: determining a normal vector associated with the second video, and rotating and translating at least one of the first videos based on the normal vector, and projecting the first video into the second video may include adding the first video to the plane. The first video may be of a first participant in the web-based video call, and the second video may be a real-world video. The plane may be a transparent two-dimensional virtual structure located in the second video. The plane may have a size proportional to the display of the device rendering the web-based video call webpage. The web-based video call webpage may include code configured to implement a trained machine learning model, and the web-based video call webpage may include code configured to implement web-based augmented reality tools. Web-based video call web pages and the web-based video call itself can be implemented with zero web application installation.

[0115] In yet another general aspect, an apparatus, system, non-transitory computer-readable medium (on which computer-executable program code capable of being executed on a computer system) and / or method can perform a process by means of: opening a web-based video call in a browser on a first device; receiving a request to join the web-based video call from a second device; capturing a first video including frames by the first device; segmenting the frames by the first device; selecting at least one segment of the segmented frames by the first device; streaming the first video including the at least one segment directly from the first device to the second device as a first real-time virtual teleport image; receiving the streamed video directly from the second device by the first device as a second video, the second video including a second real-time virtual teleport image; capturing a third video by the first device; orienting the second video based on the third video by the first device; projecting the second video onto the third video by the first device to generate a fourth video including the second real-time virtual teleport image; and rendering a webpage including the fourth video by the first device.

[0116] The implementation may include one or more of the following features. For example, initiating the web-based video call may include loading a webpage, the webpage may include code configured to implement a trained machine learning model, the trained machine learning model may be configured to segment the frames and select the at least one segment, and the webpage may include code configured to implement a web-based augmented reality tool. The web-based video call may be implemented as a zero-installation web application. The segmentation of the frames may include identifying each object in the frames, selecting the at least one segment may include selecting an object as the at least one segment, and the object may be a participant in the web-based video call. The method may further include generating a plane and positioning the plane in the second video, and the orientation of the first video may include: determining a normal vector associated with the second video, and rotating and translating at least one of the first videos based on the normal vector, and projecting the first video onto the second video may include adding the first video to the plane.

[0117] While exemplary embodiments may include various modifications and alternatives, embodiments thereof are shown by way of example in the accompanying drawings and will be described in detail herein. However, it should be understood that exemplary embodiments are not intended to be limited to the specific forms disclosed, but rather, exemplary embodiments will cover all modifications, equivalents, and alternatives falling within the scope of the claims. Throughout the description of the drawings, the same reference numerals refer to the same elements.

[0118] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuits, integrated circuits, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs executable and / or interpretable on a programmable system, which includes at least one programmable processor, which may be dedicated or general-purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to send data and instructions to the storage system, at least one input device, and at least one output device. The various implementations of the systems and techniques described herein can be implemented as and / or are generally referred to herein as circuits, modules, blocks, or systems capable of combining software and hardware aspects. For example, a module can include functional / action / computer program instructions that execute on a processor (e.g., a processor formed on a silicon substrate, GaAs substrate, etc.) or some other programmable data processing means.

[0119] Some of the exemplary embodiments described above are depicted as processes or methods illustrated as flowcharts. Although flowcharts describe operations as sequential processes, many operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of operations can be rearranged. These processes may terminate upon completion of their operations, but may also have additional steps not included in the diagram. A process can correspond to a method, function, procedure, subroutine, subroutine, etc.

[0120] The methods described above (some of which are illustrated in flowcharts) can be implemented using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments used to perform the necessary tasks can be stored in a machine or computer-readable medium, such as a storage medium. The processor can then perform the necessary tasks.

[0121] The specific structural and functional details disclosed herein are for the purpose of describing exemplary embodiments only. However, exemplary embodiments are embodied in many alternative forms and should not be construed as being limited to the embodiments set forth herein.

[0122] It should be understood that although the terms first, second, etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0123] It should be understood that when an element is described as being connected to or coupled to another element, it is either directly connected to or coupled to the other element, or there may be intermediate elements. Conversely, when an element is described as being directly connected to or directly coupled to another element, there are no intermediate elements. Other terms used to describe relationships between elements should be interpreted in a similar manner (e.g., between vs. directly between, adjacent vs. directly adjacent, etc.).

[0124] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that, when used herein, the terms *comprises*, *comprising*, *includes*, and / or *including* specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0125] It should also be noted that in some alternative implementations, the functions / actions mentioned may not occur in the order shown in the figures. For example, two figures shown consecutively may actually be executed simultaneously, or sometimes in reverse order, depending on the functions / actions involved.

[0126] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the exemplary embodiments pertain. It will be further understood that terms (e.g., those defined in common dictionaries) shall be interpreted as having a meaning consistent with their meaning in the context of the relevant field and shall not be interpreted in an idealized or overly formal sense unless expressly defined herein.

[0127] The exemplary embodiments described above, along with their corresponding detailed descriptions, are presented in the form of software or algorithms and symbolic representations of operations on data bits within computer memory. These descriptions and representations are intended to effectively convey the essence of their work to others of ordinary skill in the art. An algorithm (as used herein and in common usage) is considered a self-consistent sequence of steps that leads to a desired result. These steps are those that require physical manipulation of physical quantities. Typically, though not essential, these quantities take the form of optical, electrical, or magnetic signals that can be stored, transmitted, combined, compared, and otherwise manipulated. Primarily for general reasons, it has proven convenient to sometimes refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc.

[0128] In the illustrative embodiments described above, references to actions and symbolic representations (e.g., in the form of flowcharts) of operations that can be implemented as program modules or functional processes include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types, and can be described and / or implemented using existing hardware at existing structural elements. Such existing hardware may include one or more central processing units (CPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), computers, etc.

[0129] However, it should be remembered that all these and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to those quantities. Unless otherwise specified, or as is apparent from the discussion, terms such as processing or calculating or determining display refer to the actions and processes of a computer system or similar electronic computing device that manipulate and transform data representing physical, electronic quantities within the registers and memories of the computer system into other data representing physical quantities similarly represented within the computer system's memory or registers or other such information storage, transmission, or display devices.

[0130] It should also be noted that the software implementation aspects of the exemplary embodiments are typically encoded on some form of non-transitory program storage medium or implemented on some type of transmission medium. The program storage medium may be magnetic (e.g., floppy disk or hard disk drive) or optical (e.g., optical disc read-only memory or CD-ROM), and may be read-only or random access. Similarly, the transmission medium may be twisted-pair cables, coaxial cables, optical fibers, or some other suitable transmission media known in the art. The exemplary embodiments are not limited to these aspects of any given implementation.

[0131] Finally, it should be noted that although the appended claims set forth a particular combination of features described herein, the scope of this disclosure is not limited to the particular combination claimed herein, but extends to cover any combination of features or embodiments disclosed herein, regardless of whether that particular combination is specifically enumerated in the appended claims at this time.

Claims

1. A method for streaming video in a web-based video conferencing environment, the method comprising: Open a web-based video call in the browser on the first device; The first device receives a request from the second device to join the web-based video call; The first video, comprising frames, is captured by the first device; The frame is segmented by the first device; The first device selects at least one segment of the segmented frame; The first video, including at least one of the segments, is streamed directly from the first device to the second device as a real-time virtual teletransmission; The first device receives streaming video directly from the second device as the second video, and the second video includes a second real-time virtual teletransmitted image; A plane is generated and positioned within the second video, the plane having a size proportional to the display of the device rendering the web-based video call; The first device orients the second video based on its environment; The first device projects the second video onto a background representing the environment of the first device to generate a third video including a second real-time virtual teleportation image; as well as The third video is rendered by the first device.

2. The method according to claim 1, wherein, Establishing the web-based video call includes loading a webpage. The webpage includes code configured to implement a trained machine learning model, and The trained machine learning model is configured to segment the frame and select at least one segment.

3. The method according to claim 1, wherein, The at least one segment is an image of a participant in the web-based video call.

4. The method according to claim 1, wherein, The web-based video call is implemented using web-based communication standards.

5. The method according to claim 1, wherein, The segmentation of the frame includes: The pixels in the frame are grouped into semantic regions to locate objects and boundaries. The pixels of the frame are classified into two categories: 1) pixels representing people and 2) pixels representing the background. The pixels representing the person are segmented from the frame.

6. The method according to claim 1, wherein: The segmentation of the frame includes identifying each object in the frame. Selecting at least one fragment includes selecting an object as said at least one fragment, and The object is a participant in the web-based video call.

7. The method according to claim 1, wherein, The at least one segment is an image of a participant in the web-based video call, and the method further includes: The image is converted from a two-dimensional image to a three-dimensional image.

8. The method according to claim 1, wherein, The at least one segment is an image of a participant in the web-based video call, and the method further includes: Apply the filter to the image.

9. The method according to any one of claims 1 to 8, wherein, The web-based video call is implemented with zero web application installation.

10. The method according to claim 2, wherein, The webpage containing the code is configured to implement a web-based augmented reality tool.

11. The method according to claim 1, wherein, The second video includes: Determine the normal vector associated with the second video; Rotate and translate at least one of the first video based on the normal vector; and Projecting the first video onto the second video includes adding the first video to the plane.

12. A method for streaming video in a web-based video conferencing environment, the method comprising: Open a web-based video call in the browser on the first device; A request to join the web-based video call is transmitted from the second device; The first device directly receives the streaming video from the second device as the first video. Generate a plane having a size proportional to the display of the device rendering the web-based video call; The second video is captured by the first device; The first device directs the first video based on the second video; The first device projects the first video onto the second video to generate a third video; as well as The third video is rendered by the first device.

13. The method according to claim 12, wherein, The first video includes: Determine the normal vector associated with the second video; and Rotate and translate at least one of the first videos based on the normal vector.

14. The method of claim 12 or 13, further comprising generating a plane and positioning the plane within the second video, wherein, Projecting the first video onto the second video includes adding the first video to the plane.

15. The method of claim 12, further comprising: Generate a plane and position the plane within the second video; The first video includes: Determine the normal vector associated with the second video; and Rotate and translate at least one of the first video based on the normal vector; and Projecting the first video onto the second video includes adding the first video to the plane.

16. The method according to claim 12, wherein, The first video is a video of the first participant in the web-based video call, and the second video is a real-world video.

17. The method according to claim 12, wherein, The web-based video call includes code configured to implement a trained machine learning model, and the web-based video call includes code configured to implement a web-based augmented reality tool.

18. The method according to claim 12, wherein, The web-based video call webpage and the web-based video call are implemented with zero web application installation.

19. A method for streaming video in a web-based video conferencing environment, the method comprising: Open a web-based video call in the browser on the first device; Receive a request to join the web-based video call from the second device; A first video, comprising frames, is captured by a first device; The frame is segmented by the first device; The first device selects at least one segment of the segmented frame; The first video, including at least one of the segments, is streamed directly from the first device to the second device as a first real-time virtual teleport image; The first device receives streaming video directly from the second device as the second video, and the second video includes a second real-time virtual teletransmitted image; A plane is generated and positioned within the second video, the plane having a size proportional to the display of the device rendering the web-based video call; The third video is captured by the first device; The first device directs the second video based on the third video; The first device projects the second video onto the third video to generate a fourth video that includes the second real-time virtual teleported image; as well as The first device renders a webpage that includes the fourth video.

20. The method according to claim 19, wherein, Establishing the web-based video call includes loading a webpage. The webpage includes code configured to implement a trained machine learning model. The trained machine learning model is configured to segment the frames and select at least one segment, and The webpage includes code configured to implement web-based augmented reality tools.

21. The method according to claim 19, wherein, The web-based video call is implemented with zero web application installation.

22. The method according to claim 19, wherein, The segmentation of the frame includes: The pixels in the frame are grouped into semantic regions to locate objects and boundaries. The pixels of the frame are classified into two categories: 1) pixels representing people and 2) pixels representing the background. The pixels representing the person are segmented from the frame.

23. The method of claim 19, wherein The segmentation of the frame includes identifying each object in the frame. Selecting at least one fragment includes selecting an object as said at least one fragment, and The object is a participant in the web-based video call.

24. The method according to any one of claims 19 to 23, wherein, The orientation of the first video includes: Determine the normal vector associated with the second video, and Based on the rotation and translation of at least one of the first video vectors, and Projecting the first video onto the second video includes adding the first video to the plane.

25. An apparatus for streaming video in a web-based video conferencing environment, the apparatus comprising one or more processors and a memory storing instructions, the instructions, when executed by the one or more processors, causing the one or more processors to perform the method according to any one of claims 1 to 24.