Methods, apparatus, devices, and storage media for viewport processing

By adjusting the field of view (FoV) to match the speed of viewport changes, the motion sickness problem caused by visual dissonance during immersive video and teleconferences has been resolved, improving the user experience.

CN115552358BActive Publication Date: 2026-05-26TENCENT AMERICA LLC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2022-03-24
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In immersive video and telephone conferences, users are prone to motion sickness when following another user's viewport because the eyes cannot coordinate movement with visual information, leading to discomfort.

Method used

By adjusting the field of view (FoV) of the first user terminal and increasing or decreasing the FoV of the first user terminal according to the FoV speed of the second user terminal, the speed of the viewport change can be matched to reduce the impact of motion sickness.

Benefits of technology

It effectively reduces motion sickness when users follow another user's viewport, improving visual comfort and immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115552358B_ABST
    Figure CN115552358B_ABST
Patent Text Reader

Abstract

A method, apparatus, device, and storage medium for viewport processing are provided. The method includes: determining a first FoV based on the velocity of a second field of view (FoV) of a second user terminal, wherein the second FoV of the second user terminal is an unreduced original FoV; generating a modified first FoV by at least one of the following: (1) reducing the first FoV based on an increase in the velocity of the second FoV, and (2) increasing the first FoV based on a decrease in the velocity of the second FoV; and transmitting the modified first FoV to a first user terminal and presenting the modified first FoV as a new viewport on the first user terminal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application is based on and claims priority to U.S. Provisional Patent Application No. 63 / 167,304, filed March 29, 2021, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] Embodiments of this disclosure relate to methods, apparatus, devices, and storage media for viewport processing to reduce motion sickness in teleconferences and telepresentations of remote terminals, and more specifically to reducing motion sickness in a user when the user follows another user's viewport. Background Technology

[0004] When using omnidirectional media streaming, only the portion of the content corresponding to the user's viewport is presented, while a head-mounted display (HMD) is used to give the user a realistic view of the media stream.

[0005] Figure 1 The illustration depicts a relevant scenario (Scenario 1) for an immersive teleconference call, where the call is organized between room A (101), user B (102), and user C (103). Figure 1 As shown, room A (101) represents a conference room with an omnidirectional / 360-degree camera (104), and users B (102) and C (103) are remote participants using an HMD and a mobile device, respectively. In this scenario, participants B (102) and C (103) send their viewport orientations to room A (101), and room A (101) in turn sends viewport-related streams to users B (102) and C (103).

[0006] Figure 2A The diagram illustrates an extended scenario (Scenario 2) comprising multiple meeting rooms (2a01, 2a02, 2a03, 2a04). User B (2a06) uses an HMD to view a video stream from a 360-degree camera (104), and User C (2a07) uses a mobile device to view the video stream. Users B (2a06) and C (2a07) send their viewport orientations to at least one of the meeting rooms (2a01, 2a02, 2a03, 2a04), and at least one of the meeting rooms (2a01, 2a02, 2a03, 2a04) in turn sends viewport-related streams to users B (2a06) and C (2a07).

[0007] like Figure 2BAs shown, another example scenario (Scenario 3) is when a call is established using an MRF / MCU (2b05), where the Media Resource Function (MRF) and Media Control Unit (MCU) are multimedia servers that provide media-related functions for bridging terminals in a multi-party conference call. Conference rooms can send their respective videos to the MRF / MCU (2b05). These videos are viewport-independent; that is, the entire 360-degree video is sent to the media server (i.e., the MRF / MCU), independent of the user's viewport for streaming a specific video. The media server receives the viewport orientation of users (User B (2b06) and User C (2b07)) and sends viewport-related streams to the users accordingly.

[0008] Further for scenario 3, remote users can choose to view one of the available 360-degree videos from the conference rooms (2a01 to 2a04, 2b01 to 2b04). In this case, the user sends the video to be streamed and its viewport orientation information to the conference room or the MRF / MCU (2b05). Users can also switch from one room to another based on active speaker triggering.

[0009] Another extension of the above scenario is, for example, when user A wearing the HMD is interested in the viewport of another user. In this particular example, the other user would be user B (102). This could happen when user B (102) is presenting to the conference room (room A, 2a01, 2a02, 2a03, and / or 2a04) or when user A is interested in user B (102)'s focus or viewport. However, when this happens and user A's viewport is switched, user A may experience motion sickness. Summary of the Invention

[0010] One or more exemplary embodiments of this disclosure provide a method, apparatus, device, and storage medium for viewport processing.

[0011] According to an embodiment, a method for viewport processing is provided. The method may include: determining a first FoV based on the velocity of a second field of view (FoV) of a second user terminal, wherein the second FoV of the second user terminal is an unreduced original FoV; generating a modified first FoV by at least one of the following: (1) reducing the first FoV based on an increase in the velocity of the second FoV, and (2) increasing the first FoV based on a decrease in the velocity of the second FoV; and transmitting the modified first FoV to a first user terminal and presenting the modified first FoV as a new viewport on the first user terminal.

[0012] According to an embodiment, an apparatus is provided. The apparatus may include one or more memories configured to store program code, and one or more processors. When the program code is executed by the at least one processor, the at least one processor performs a method according to an embodiment of this application. According to an embodiment, an apparatus for viewport processing is also provided, the apparatus comprising: a determining module for determining a first FoV based on the speed of a second field of view (FoV) of a second user terminal, wherein the second FoV of the second user terminal is an unreduced original FoV; a modifying module for generating a modified first FoV by at least one of the following: (1) reducing the first FoV based on an increase in the speed of the second FoV, and (2) increasing the first FoV based on a decrease in the speed of the second FoV; and a transmitting module for transmitting the modified first FoV to a first user terminal and presenting the modified first FoV as a new viewport on the first user terminal.

[0013] According to an embodiment, a non-volatile computer-readable medium is provided for alleviating motion sickness when a first user follows the viewport of a second user. The computer-readable medium may be connected to one or more processors and may be configured to store instructions that, when executed by at least one processor of a device, cause at least one or more processors to perform a method according to an embodiment of this application.

[0014] Additional aspects will be set forth in part in the description which follows, and will be apparent in part from the description, or may be learned by practice of the embodiments of this disclosure presented. Attached Figure Description

[0015] The above and other aspects, features and aspects of embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings.

[0016] Figure 1 This is a diagram illustrating the ecosystem used for immersive teleconferencing.

[0017] Figure 2A This is a diagram of a multi-party, multi-conference telephone conference.

[0018] Figure 2B This is a diagram illustrating a multi-party, multi-conference room teleconference using MRF / MCU.

[0019] Figure 3 It is a simplified block diagram of a communication system according to one or more embodiments.

[0020] Figure 4 This is a simplified example illustration of a streaming environment according to one or more embodiments.

[0021] Figure 5 This is a schematic diagram of the field of view (FoV) change according to an embodiment.

[0022] Figure 6 This is a flowchart of a method for alleviating motion sickness when a user follows another user's viewport in a streaming session, according to an embodiment.

[0023] Figure 7 This is a schematic diagram of a computer system according to an embodiment. Detailed Implementation

[0024] This disclosure relates to a method and apparatus for alleviating motion sickness when a user follows another user's viewport.

[0025] Motion sickness occurs when the human eye cannot coordinate perceived movement with actual body movement. The brain becomes confused because it receives conflicting information from the balance organs in the inner ear and eyes. This phenomenon is common in immersive video and affects immersive video streaming and immersive teleconferencing. Figure 2A and Figure 2B As shown, multiple conference rooms equipped with omnidirectional cameras are in the middle of a teleconference. Users can select the video stream to be displayed as an immersive stream from one of the conference rooms (2a01, 2a02, 2a03, 2a04) or from the viewport of another user participating in the teleconference.

[0026] Embodiments of this disclosure will be fully described with reference to the accompanying drawings. However, examples of embodiments may be implemented in a variety of forms, and this disclosure should not be construed as limited to the examples described herein. Rather, examples of embodiments are provided to make the technical solutions of this disclosure more comprehensive and complete, and to fully convey the ideas of the examples of embodiments to those skilled in the art. The drawings are merely illustrative examples of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of these parts are omitted.

[0027] The features discussed below can be used individually or in any combination in any order. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. Furthermore, these embodiments can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits), or in software, or in different network and / or processor devices and / or microcontroller devices. In one example, one or more processors execute a program stored on a non-volatile computer-readable medium.

[0028] Figure 3This is a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) may include at least two terminals interconnected via a network (305), namely a second terminal (302) and a third terminal (303). For unidirectional data transmission, the third terminal (303) may encode video data at a local location for transmission to the second terminal (302) via the network (305). The second terminal (302) may receive the encoded video data from the third terminal from the network (305), decode the encoded data, and display the recovered video data. Unidirectional data transmission is common in media service applications such as teleconferencing.

[0029] Figure 3 The illustration shows a second pair of terminals, namely a first terminal (301) and a fourth terminal (304), which are provided to support bidirectional transmission of encoded video, for example, during video conferencing. For bidirectional data transmission, each of the first terminal (301) and the fourth terminal (304) can encode video data captured at a local location for transmission to the other terminal via a network (305). Each of the first terminal (301) and the fourth terminal (304) can also receive encoded video data transmitted by the other terminal and display the recovered video data on a local display device.

[0030] exist Figure 3 In this application, the first terminal (301), the second terminal (302), the third terminal (303), and the fourth terminal (304) can be exemplified as servers, personal computers, and smartphones, but the principles disclosed herein are not limited thereto. The embodiments disclosed herein are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (250) refers to any number of networks, including, for example, wired (connected) and / or wireless communication networks, that transmit encoded video data between the first terminal (301), the second terminal (302), the third terminal (303), and the fourth terminal (304). The communication network (250) can exchange data in circuit-switched and / or packet-switched channels. This network may include telecommunications networks, local area networks, wide area networks, and / or the Internet. The immersive video discussed in the embodiments of this disclosure can be sent and / or received via networks (305), etc.

[0031] Figure 4 The illustration shows an example streaming environment used for the disclosed topics. The disclosed topics can be equivalently applied to other video-enabled applications, including, for example, immersive teleconferencing, video teleconference, and telepresence.

[0032] The streaming environment may include one or more meeting rooms (403), which may include video sources (401), such as video cameras, and one or more participants in the meeting (402). Figure 4 The video source (401) illustrated in the figure is, for example, a 360-degree video camera capable of creating a video sample stream. The video sample stream can be sent to and / or stored on the streaming server (404) for future use. One or more streaming clients (405, 406) can also send their respective viewport information to the streaming server (404). Based on the viewport information, the streaming server (404) can send the viewport-related stream to the corresponding streaming client (405, 406). In another example embodiment, the streaming client (405, 406) can access the streaming server (404) to retrieve the viewport-related stream. The embodiments are not limited to this configuration; one or more conference rooms (403) can communicate with the streaming clients (405, 406) via a network (e.g., network 305). Additionally, the streaming server (404) and / or the streaming clients (405, 406) can include hardware, software, or a combination thereof to allow or implement aspects of the disclosed subject matter, which are described in more detail below. The streaming client (405, 406) may include FoV (Field of View) components (407a, 407b). The FoV components (407a, 407b) may adjust the viewport or FoV of the streaming client according to embodiments described in more detail below, and create an output video sample stream that can be reproduced on the display 408 or other playback devices (e.g., HDM, speakers, mobile devices, etc.).

[0033] In an immersive teleconference call, the streaming client (hereinafter referred to as the “user”) can choose to switch between one of the available immersive videos from multiple rooms (e.g., one or more conference rooms (403)) from which immersive 360-degree video is streamed, or there is no available immersive video. Immersive video can be switched automatically or manually from one room to another. When a user manually switches immersive video from one source to another, the user is prepared for the switch. Therefore, the chance of the user experiencing motion sickness can be reduced. However, when the video is switched automatically, the user may not be prepared for the switch and may experience motion sickness. For example, when user A’s viewport is automatically switched to follow another user B’s viewport. When this happens, user A’s eyes cannot coordinate the movement in user B’s viewport with the information received by user A’s inner ear and / or eyes.

[0034] In some embodiments, when user A follows the viewport of another user B, user A's FoV is reduced to mitigate motion sickness and provide greater visual comfort. FoV can be defined as a function of the HMD speed / dynamics of the FoV of the following user. Therefore, user A's FoV can be adjusted accordingly to effectively reduce the effects of motion sickness caused by user A following the viewport of another user B. For example, user A's FoV can be reduced when the HMD speed / dynamics of user B's (followed by user A) FoV increases. In the same or another example embodiment, user A's FoV can be increased when the HMD speed / dynamics of user B's (followed by user A) FoV decreases.

[0035] Figure 5 The diagram illustrates the changes in FoV of user A following another user B, where the HMD velocity / dynamics of the FoV of the other user B is increasing.

[0036] like Figure 5 As shown, images (50a, 50b, 50c) are illustrated as different levels of HMD velocity / dynamics with FoV of the followed user B. Each image (50a, 50b, and 50c) contains high-resolution portions (501, 503, 505) and low-resolution portions (502, 504, 506). The corresponding high-resolution portions (501, 503, 505) and low-resolution portions (502, 504, 506) of images (50a, 50b, 50c) represent the HMD velocity / dynamics based on the FoV of user B. Figure 5 The HMD speed / dynamics of the FoV in image 50a is less than that in image 50b, and the HMD speed / dynamics of the FoV in image 50b is less than that in image 50c (i.e., HMD speed in image 50a < HMD speed in image 50b < HMD speed in image 50c). When user A's viewport changes to user B's viewport, the high-resolution portions (501, 503, 505) of images (50a, 50b, 50c) can be reduced according to the increase in the HMD speed / dynamics of the FoV to mitigate the effects of motion sickness. Therefore, the high-resolution portion 501 in image 50a is greater than the high-resolution portion 503 in image 50b, and the high-resolution portion 503 is greater than the high-resolution portion 505 in image 50c (i.e., high-resolution portion 501 > high-resolution portion 503 > high-resolution portion 505). The high-resolution portion of the image increases as the HMD speed / dynamics decreases. Similarly, the low-resolution portion of the image shrinks as the speed / dynamics of the HMD decreases.

[0037] In the same or another embodiment, when user A's FoV is reduced, the initial FoV (without reduction) and the reduced FoV (hereinafter referred to as "FoV") are... Reduced The area between “” and “) can be transmitted at low resolution. For example, refer to Figure 5 In image 50a, when user A's FoV is reduced to the high-resolution portion 501, the initial FoV (i.e., image 50a) and FoV without reduction are... Reduced The region between (i.e., the high-resolution portion 501) and (i.e., the low-resolution portion 502) can be transmitted at low resolution.

[0038] In the same or another embodiment, when user A's FoV is reduced, only the FoV can be reduced at high resolution, low resolution, or a combination thereof. Reduced Transmitted to user A.

[0039] In the same or another embodiment, a reduction factor λ can be defined. The value of the reduction factor λ can be applied to the initial or original FoV to obtain the FoV. Reduced (For example, high-resolution sections 501, 503, and / or 505), the FoV Reduced The aim is to ensure better visual comfort for users who are changing their view (e.g., user A changing their view to that of another user B). FoV Reduced The relationship between the reduction factor λ and the reduction factor can be described by the following equation:

[0040] FoV Reduced =λFoV / HMD Speed ​​(1)

[0041] In the same or another embodiment, a minimum value of reduced FoV (hereinafter referred to as "FoV") can be defined for user A. min To avoid compromising immersion, FoV should be reduced excessively. For example, excessively reducing FoV will affect user acceptability and interfere with the video's immersive experience. Therefore, FoV... min Less than or equal to FoV Reduced As described below:

[0042] FoV min <=FoV Reduced (2)

[0043] In the same or another embodiment, the reduction factor λ can have values ​​between 1 and (FoV). min The reduction factor λ is a value between λ and λ (FoV), and the weight decreases linearly between the ends. A reduction factor λ with a value of 1 means that the client is prone to motion sickness. The value of the reduction factor λ can be set by the user following the viewport (e.g., user A). Accordingly, the reduction factor λ provides the receiving user with control over the reduction of their FoV.

[0044] According to the embodiment, before receiving the FoV from user B, user A needs to input the value of the reduction factor λ and the FoV. min This is transmitted to the sender (User B) as part of its device capabilities, allowing the reduced FoV to be sent to User A in an optimal manner (e.g., at the necessary bit rate). This can be sent at the start of the session or during the session, for example, via SDP required by User A. Sending the reduced high-resolution FoV by the sender (User B) reduces bandwidth requirements compared to the receiving user reducing the FoV after receiving the original FoV.

[0045] Figure 6 This is a flowchart of a method 600 for viewport processing according to an embodiment. Method 600 may be executed, for example, by an electronic device. The electronic device may be, for example, a user terminal such as a streaming client.

[0046] like Figure 6 As shown, in step S610, method 600 includes determining whether the first user terminal changes its viewport to the viewport of the second user terminal. If the result is no at step S610, method 600 repeats step S610. If the result is yes at step S610, method 600 continues to step S620.

[0047] In step S620, method 600 determines the first FoV of the first user terminal based on the speed of the second FoV of the second user terminal. When the speed of the second FoV of the second user terminal changes, the first FoV of the first user terminal is determined and / or modified. The second FoV of the second user terminal is the original FoV that has not been reduced or increased.

[0048] In step S630, method 600 determines whether the second speed of the FoV of the second user terminal has increased. If yes at step S630, the first FoV of the first user terminal is reduced (S640), and the reduced first FoV is transmitted to the first user terminal (S670). If no at step S630, method 600 returns to step S620.

[0049] In step S650, method 600 determines whether the speed of the second FoV of the second user terminal has decreased. If yes at step S650, the first FoV of the first user terminal is increased (S660), and the increased first FoV is transmitted to the first user terminal (S670). If no at step S650, method 600 returns to step S620. Based on step S670, it is known that the embodiments of this application can transmit the modified first FoV to the first user terminal so that the modified first FoV is presented as a new viewport on the first user terminal. In this way, when the first user follows the viewport of the second user in a streaming session, method 600 helps to alleviate motion sickness in the user.

[0050] Although Figure 6 An example box of the method is shown, but in some implementations, the method may include... Figure 6 The boxes depicted in the diagram may be fewer, different, or arranged differently compared to additional boxes. Alternatively, the method may be applied to two or more boxes in parallel.

[0051] The techniques described above for mitigating motion sickness in immersive teleconferencing and remote presentation can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 7 A computer system 700 suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0052] Computer software can be coded using any suitable machine code or computer language. Machine code or computer language can be created through assembly, compilation, linking or similar mechanisms. This code includes instructions that can be executed directly by a computer's central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode execution, or other means.

[0053] These instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, and Internet of Things (IoT) devices.

[0054] Figure 7 The components shown for the computer system 700 are exemplary in nature and are not intended to impose any limitation on the scope or functionality of the computer software used to implement the embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement on any component or combination thereof illustrated in the exemplary embodiments of the computer system 700.

[0055] Computer system 700 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users via, for example, tactile input (such as keystrokes, swipes, data glove movements), audio input (such as sound, tapping), visual input (such as gestures), or olfactory input. The human-machine interface device may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as speech, music, ambient sound), images (such as scanned images, photographic images obtained from still image cameras), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).

[0056] The input human-machine interface device may include one or more of the following (only one of each is depicted): keyboard 701, trackpad 702, mouse 703, touch screen 709, data glove, joystick 704, microphone 705, camera 706, scanner 707.

[0057] The computer system 700 may also include certain human-machine interface (HMI) output devices. These HMI output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such HMI output devices may include tactile output devices (e.g., tactile feedback via a touchscreen 709, data gloves, or joystick 704, but may also include tactile feedback devices that are not used as input devices), audio output devices (such as speakers 708, headphones), visual output devices (such as screens 709, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability—some of which may be capable of outputting two-dimensional or more than three-dimensional visual output through means such as stereoscopic output; virtual reality glasses, holographic displays, and smoke canisters), and printers.

[0058] The computer system 700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 711 with media 710 such as CD / DVD, thumb drives 712, removable hard disk drives or solid-state drives 713, conventional magnetic media such as magnetic tapes and floppy disks, and devices based on dedicated ROM / ASIC / PLD such as SecureDell chips.

[0059] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other volatile signals.

[0060] Computer system 700 may also include an interface 715 to one or more communication networks 714. Network 714 may be, for example, wireless, wired, or optical. Network 714 may also be local, wide area, metropolitan, vehicular, industrial, real-time, latency-tolerant, etc. Examples of network 714 include local area networks (such as Ethernet), wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicle and industrial networks (including CANbus), etc. Some networks 714 typically require an external network interface adapter (e.g., graphics adapter 725) attached to certain general-purpose data ports or peripheral buses 716 (such as, for example, a USB port of computer system 700); other networks are typically integrated into the core of computer system 700 by attaching to system buses as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smartphone computer system). Using any of these networks 714, computer system 700 can communicate with other entities. This communication can be unidirectional, receive-only (e.g., broadcasting TV), transmit-only (e.g., CANbus to certain CANbus devices), or bidirectional (e.g., to other computer systems using a local area digital network or a wide area digital network). Certain protocols and protocol stacks can be used on each of the networks and network interfaces described above.

[0061] The aforementioned human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the kernel 717 of the computer system 700.

[0062] The core 717 may include one or more central processing units (CPUs) 718, graphics processing units (GPUs) 719, dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) 720, hardware accelerators 721 for certain tasks, etc. These devices, along with read-only memory (ROM) 723, random access memory (RAM) 724, and internal mass storage devices 722 such as internal non-user-accessible hard disk drives (SDs), etc., can be connected via a system bus 726. In some computer systems, the system bus 726 may be accessed as one or more physical connectors to allow for expansion by additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus 726 or via a peripheral bus 716. Peripheral bus architectures include PCI, USB, etc.

[0063] The CPU 718, GPU 719, FPGA 720, and accelerator 721 can execute certain instructions, and combinations of these instructions can constitute the aforementioned computer code. This computer code can be stored in ROM 723 or RAM 724. Transient data can also be stored in RAM 724, while permanent data can be stored, for example, in an internal mass storage device 722. Fast storage and retrieval of any memory device can be enabled by using a cache memory, which can be closely associated with one or more CPUs 718, GPUs 719, mass storage devices 722, ROM 723, RAM 724, etc.

[0064] Computer-readable media may contain computer code for performing operations of various computer implementations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be of types known and available to those skilled in the art of computer software.

[0065] By way of example and not limitation, a computer system having architecture 700, and particularly kernel 717, can provide functionality as a result of executing software embodied in one or more tangible computer-readable media (including CPUs, GPUs, FPGAs, accelerators, etc.). Such computer-readable media can be media associated with user-accessible mass storage devices as described above, as well as certain storage devices of kernel 717 having non-volatile properties (such as internal kernel mass storage 722 or ROM 723). Software implementing various embodiments of this disclosure can be stored in such devices and executed by kernel 717. Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause kernel 717, and particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM 724 and modifying such data structures according to software-defined processes. Alternatively or as an alternative, the computer system may provide functionality (e.g., accelerator 721) as a result of logical hardwired connections or otherwise embodied in circuitry, which may replace or operate with software to perform the specific process or a specific portion of the specific process described herein. References to software may, where appropriate, include logic, and vice versa. References to computer-readable media may, where appropriate, include circuitry storing software for execution (such as integrated circuits (ICs)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0066] While several exemplary embodiments have been described in this disclosure, there are changes, substitutions, and various alternative equivalents that fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and are thus within its spirit and scope.

Claims

1. A method for viewport processing, characterized in that, The method includes: The first FoV is determined based on the velocity of the second field of view (FoV) of the second user terminal, wherein the second FoV of the second user terminal is the original FoV without reduction, the first FoV represents the high-resolution region of the image presented to the first user terminal, and the region between the first FoV and the second FoV represents the low-resolution region of the image presented to the first user terminal. The modified first FoV is generated by at least one of the following methods: (1) reducing the first FoV based on the velocity increase of the second FoV, and (2) increasing the first FoV based on the velocity decrease of the second FoV; and The modified first FoV is adjusted by applying a reduction factor to the second FoV, wherein the reduction factor is defined by the first user terminal, and the adjusted first FoV is equal to the reduction factor multiplied by the second FoV and then divided by the velocity of the second field of view FoV; The minimum FoV is transmitted to the second user terminal via the Session Description Protocol (SDP), wherein the minimum FoV is defined by the first user terminal, and the first FoV of the first user terminal is not less than the minimum FoV; The modified first FoV is transmitted to the first user terminal, and the modified first FoV is presented as a new viewport on the first user terminal.

2. The method of claim 1, further comprising determining the speed of the second FoV as the speed of a head-mounted display worn by the second user or the speed of a handheld device operated by the second user.

3. The method of claim 1, further comprising transmitting the image region between the second FoV and the modified first FoV at a low resolution based on the reduction of the first FoV.

4. The method according to claim 1, wherein, The method further includes: when the first FoV is reduced, transmitting only the image region corresponding to the reduced first FoV to the first user terminal at high resolution, low resolution, or a combination of both.

5. The method according to claim 1, wherein, The method further includes: determining whether the first user terminal follows the viewport of the second user terminal in a streaming session; Specifically, when determining that the first user terminal follows the viewport of the second user terminal, the speed of the second field of view (FoV) of the second user terminal is used to determine the first FoV.

6. The method according to claim 1 further includes transmitting the reduction factor to the second user terminal via Session Description Protocol (SDP).

7. The method according to claim 6, wherein, The reduction factor is transmitted to the second user terminal at the start of the streaming session or during the streaming session.

8. The method according to claim 1, wherein, The minimum FoV is transmitted to the second user terminal at the start of the streaming session or during the streaming session.

9. A device, characterized in that, The device includes: At least one memory is configured to store program code; and At least one processor; When the program code is executed by the at least one processor, the at least one processor performs the method as described in any one of claims 1-8.

10. An apparatus for viewport processing, characterized in that, The device includes: The determining module determines a first FoV based on the speed of the second field of view (FoV) of the second user terminal, wherein the second FoV of the second user terminal is the original FoV without reduction, the first FoV represents the high-resolution region of the image presented to the first user terminal, and the region between the first FoV and the second FoV represents the low-resolution region of the image presented to the first user terminal. The modification module generates a modified first FoV by at least one of the following methods: (1) reducing the first FoV based on the speed increase of the second FoV, and (2) increasing the first FoV based on the speed decrease of the second FoV; adjusting the modified first FoV by applying a reduction factor to the second FoV, wherein the reduction factor is defined by the first user terminal, and the adjusted first FoV is equal to the reduction factor multiplied by the second FoV and divided by the speed of the second field of view FoV; and The transmission module is configured to: transmit a minimum FoV to the second user terminal via the Session Description Protocol (SDP), wherein the minimum FoV is defined by the first user terminal, and the first FoV of the first user terminal is not less than the minimum FoV; transmit the modified first FoV to the first user terminal, and present the modified first FoV as a new viewport on the first user terminal.

11. A non-volatile computer-readable medium storing instructions, characterized in that, The instructions include: one or more instructions that, when executed by at least one processor, cause the at least one processor to: perform the method as described in any one of claims 1-8.