Coding rate allocation technology of low-delay video communication system

By dividing the video stream into slices in a wide-angle video communication system and dynamically adjusting the bit rate, the problem of network bandwidth and computing resources is solved, and high-quality image display and low-latency transmission are achieved.

CN120264096APending Publication Date: 2025-07-04AGORA LAB INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411380270.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-04
Filing Date
2024-09-30
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing wide-angle video communication systems are wasted on network bandwidth consumption and computing resources, and the prior art is difficult to effectively allocate code rates in low latency to improve image quality.

Method used

By dividing the video stream into slices associated with the viewing angle, and dynamically adjusting the bitrate allocation of each slice at the sending and receiving ends, the bitrate rate is adjusted according to the image quality score to optimize image quality and reduce network bandwidth consumption.

Benefits of technology

It improves the image quality of the receiver, reduces network bandwidth consumption and device processing power consumption, and improves user viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264096A_ABST
    Figure CN120264096A_ABST
Patent Text Reader

Abstract

The present invention provides a method for allocating code rates in video communications, comprising: encoding, by an encoder, a video stream, the video stream being divided into slices associated with a view of the video stream, where encoding the video stream comprises allocating a code rate for each slice; determining an image quality score for slice rendering in the first iteration, wherein the slice rendering comprises decoding and displaying slices in the first iteration in a perspective of the video stream at the allocated code rate; and after determining the image quality score, adjusting a code rate assigned to the one or more slices such that rendering of the slices in a second iteration after the first iteration includes decoding and displaying the slices with the adjusted code rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to video encoding and transmission technologies, and mainly relates to real-time video communication systems. Background Art

[0002] Video quality is an important aspect in real-time multimedia communication systems. In such systems, video data is transmitted between a sending device and a receiving device (such as a mobile phone and a personal computer) via a network (such as the Internet). Objective metrics are usually available for evaluating and controlling video quality during network transmission. For example, the objective metric can usually be (or can include) peak signal-to-noise ratio (PSNR) and structural similarity (SSIM), etc.

[0003] Video quality may be affected by various network instability factors, such as low network transmission speed, connection problems, and network "congestion", etc. That is to say, unstable network bandwidth may have a negative impact on video quality during network transmission. The video encoder may adjust its encoding scheme and parameters (such as video encoding bitrate, frame rate, and / or picture resolution) according to different network conditions, which may lead to fluctuations in video quality. Such fluctuations may ultimately affect the user experience in real-time multimedia communication systems. Summary of the Invention

[0004] On the one hand, the present invention provides a method for allocating bitrate in video communication, including: encoding a video stream by an encoder, where the video stream is divided into slices associated with the view of the video stream, and encoding the video stream includes allocating a bitrate for each slice; scoring the image quality of slice rendering at the first iteration, where slice rendering includes decoding and displaying the slice at the first iteration at the bitrate in the view of the video stream; and after determining the image quality score, modifying the bitrates allocated to one or more slices such that the slices are decoded and displayed using the adjusted bitrates when rendered at the second iteration after the first iteration.

[0005] On the other hand, the present invention provides a device for allocating bitrates in video communication, including a non-transitory memory and a processor configured to execute instructions stored in the non-transitory memory. The instructions stored in the non-transitory memory include: instructions for encoding a video stream that is divided into slices associated with the viewpoints of the video stream, wherein encoding the video stream includes allocating bitrates for each slice; instructions for scoring the image quality of slice rendering at the first iteration, wherein slice rendering includes decoding and displaying the slices at the first iteration at the bitrates in the viewpoints of the video stream; and instructions for modifying the bitrates allocated to one or more slices after determining the image quality score, such that the adjusted bitrates are used for decoding and displaying the slices when rendering the slices at the second iteration after the first iteration.

[0006] In a third aspect, the present invention provides a non-transitory computer-readable storage medium. Such a non-transitory computer-readable storage medium is configured to store a computer program for allocating bitrates in video communication. The computer program includes: computer instructions executable by a processor for encoding a video stream that is divided into slices associated with the viewpoints of the video stream, wherein encoding the video stream includes allocating bitrates for each slice; instructions executed by the processor for scoring the image quality of slice rendering at the first iteration, wherein slice rendering includes decoding and displaying the slices at the first iteration at the bitrates in the viewpoints of the video stream; and instructions executed by the processor for modifying the bitrates allocated to one or more slices after determining the image quality score, such that the adjusted bitrates are used for decoding and displaying the slices when rendering the slices at the second iteration after the first iteration. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] When reading the following detailed description, reference to the accompanying drawings will help better understand the content of the present invention. It should be noted that, as a convention, the various parts in the drawings are not drawn to actual scale, but the dimensions of the various parts are appropriately enlarged or reduced in the drawings for clarity of expression.

[0008] Figure 1 is a schematic diagram of a system for media transmission;

[0009] Figure 2A is a schematic diagram of an image slice of a video stream;

[0010] Figure 2B is a schematic diagram of an encoded slice of an image in a video stream;

[0011] Figure 3 is a schematic diagram of a network for real-time video communication;

[0012] Figure 4A is an example flowchart of a technique for allocating bitrates for video communication;

[0013] Figure 4B Yes Figure 4A Schematic diagram of an operation mode of a technique for allocating bitrates for video communication; and

[0014] Figure 5 Is an example flowchart of a technique for allocating bitrates for video communication. Detailed implementation manners

[0015] A video communication system may include a transmitting end (i.e., a transmitting device) and a receiving end (i.e., a receiving device). The transmitting end may perform at least one of a series of operations such as video capture, video distortion or stitching, video encoding, and video transmission. In one embodiment, the transmitting end may be a client device that captures and transmits (such as video stream transmission) video in real time to one or more receiving ends. In another embodiment, the transmitting end may also be a video stream server that can receive real-time or pre-recorded video and transmit it to one or more receiving end devices in the form of a video stream. The receiving end may perform operations such as video decoding, video distortion correction, and video rendering. The transmitting end and the receiving end may communicate through a network. That is to say, the encoded video data may be transmitted from the transmitting end to the receiving end through the network. The video data may be transmitted from the transmitting end to the receiving end through multiple servers on the network.

[0016] The captured video may be a wide-angle video, which may consist of a series of images arranged in chronological order. An image in the video may be composed of parts stitched together, and these parts may also be referred to as faces or sub-faces, which may be captured by multiple cameras of a wide-angle shooting device, for example. For illustration, 6 camera devices may be used to shoot a wide-angle video, where the cameras are distributed in a cubic shape, jointly forming a field of view (FOV) of up to 360°. Each camera may cover the FOV in the longitudinal dimension (such as 120°) and the FOV in the transverse dimension (such as 90°). Therefore, the FOV of any one camera may overlap with the FOV of other cameras. The overlapping area may be used for stitching operations to generate a wide-angle video.

[0017] There are already some traditional techniques for encoding and decoding wide-angle videos. For example, the scalable high efficiency video coding extension (SHVC) of the joint collaborative team on video coding (JCT-VC) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11 can be used to encode or decode wide-angle videos, or similar techniques. Of course, other techniques can also be adopted. For example, encoding the images of a wide-angle video may include the operation of dividing the images into slices. Each slice (i.e., the slices generated in time series) can be encoded into a separate compressed bitstream (referred to as "compressed slice stream" or simply "slice stream" in this article). The images of the wide-angle video can be 8K images (i.e., 7,680×4,320 pixels), which are divided into 720p slices (i.e., each slice is 1280×720 pixels), and no specific limitation is made here. Taking the H.264 nomenclature as a non-limiting example, the slice sequence can be grouped, and each group is encoded or decoded according to an encoding structure such as IBBPBB, where I represents an intra-predicted slice, P represents a slice predicted based on the previous slice, and B represents a slice predicted based on the previous and the next slices.

[0018] As mentioned above, the viewer can only view a part of the wide-angle video at any time. That is to say, the viewer may only be able to view a part of the entire imaging scene at any time. The area that the viewer can view or is interested in viewing is the "viewport". It can also be said that the viewport is within the region of interest (ROI) or included in the region of interest (ROI). Therefore, in the present invention, the ROI refers to the wide-angle video image area that contains the viewport.

[0019] When watching a video, the viewer can change the viewport at any time. For example, the viewer may use a receiving device (such as a handheld device, a head-mounted device, or any other device capable of presenting a wide-angle video) to watch a video stream of a runner passing through the scene from left to right. Therefore, when the runner moves to the right, the viewer can move the viewport to the right. In another example, when the viewer's line of sight follows the runner and the viewer hears the sound of a low-flying plane, the viewer can change the perspective by adjusting the receiving device to face upward to see the plane. Generally, a wide-angle video communication system may include a view point prediction function, which is used to predict which slices the receiving device will use in the next (e.g., after 1, 2, 3 seconds or a certain future time point).

[0020] In a traditional wide-angle video communication system, the transmitting end can encode and transmit a complete image (e.g., a complete image in ERP format), and the receiving end can receive and decode the image according to the quality of the originally captured image, even if the receiving end cannot view all this data at one time. Transmitting and receiving the entire wide-angle video in high quality consumes a large amount of bandwidth. However, such bandwidth consumption is unnecessary, that is, wasted, because as described above, the viewer cannot view the entire wide-angle video. In addition, encoding the full frame of the wide-angle video may unnecessarily consume the computing resources of the encoder (i.e., the transmitting end), thereby reducing the performance of the transmitting end, especially in real-time applications. Similarly, decoding the full frame of the wide-angle video may also unnecessarily consume the computing resources of the receiving end (i.e., the decoder), thereby reducing the viewing (display) performance of the receiving end.

[0021] Some traditional wide-angle video communication systems can use hierarchical encoding / decoding techniques to encode the wide-angle video into a base layer and one or more enhancement layers. The base layer can encode the whole of the wide-angle video (i.e., all images) in lower quality. The base layer can include the wide-angle video or a downsampled version of the wide-angle video. The enhancement layer can include different slices divided from the original wide-angle video and encoded with a higher total bit rate and resolution than the original wide-angle video. The receiving end (e.g., the decoder of the receiving device) can always decode the base layer data and then, by adaptively requesting the enhancement layer data, decode the data corresponding to the current viewport (i.e., the encoded data of the slice).

[0022] Compared with transmitting (e.g., from the transmitting end) a complete, high-quality wide-angle image, such techniques can save bandwidth. However, in such techniques, different slices can be independently encoded and transmitted, and then stitched and remapped by the receiving end for display. Due to the different image complexities and positions within the viewport among the slices, if the encoding bit rate allocation for each slice is unreasonable, the quality of the reconstructed image may be very poor. Therefore, the reconstructed image may have inconsistent image quality and obvious slice boundaries, etc.

[0023] In addition, current slice-based wide-angle transmission usually adopts a streaming-based solution, where the server can pre-store future encoded content (e.g., images). Using such a technique, the future quality rate ratio information of different slices can be obtained in advance to help effectively allocate the bit rate to each slice. However, such a technique may not have low-latency characteristics because the pre-stored future encoded content may need to be delayed by several seconds or even minutes during streaming. Therefore, such a technique may not be able to meet more stringent low-latency requirements (e.g., instant encoding and transmission, real-time applications) because the quality rate ratio of each future slice cannot be obtained in advance when determining the transmission bit rate.

[0024] In addition to improving the viewing experience at the receiving end, the present invention can also reduce the bandwidth consumption of the network and the processing power consumption of the receiving and transmitting devices. Dynamic bitrate allocation is performed at the transmitting end and / or the receiving end so as to actively allocate the bitrate to each slice according to the quality of the previous slice images. For example, the initial image of a wide-angle video can be divided into slices, and thus each slice can be assigned an initial bitrate. The transmitting end can encode the slices and send the slices to the receiving end for rendering (such as decoding and / or displaying the images). The transmitting end and / or the receiving end can evaluate the quality of the rendered images to determine whether the bitrate assigned to the slices needs to be adjusted. That is to say, according to the quality of the rendered images, the transmitting end can re-allocate the bitrate to one or more slices for future rendering (i.e., rendering subsequent images of the wide-angle video after rendering the initial image). Therefore, the image quality can be improved by re-assigning the total bitrate quota (i.e., capacity) to the slices according to the bandwidth consumption of the network and / or the processing power consumption of the receiving device and the transmitting device, thereby improving the user viewing experience of the wide-angle video.

[0025] To describe the embodiments of the present invention in more detail, first, an example of the hardware and software devices for implementing a real-time wide-angle video communication system will be understood. It should be noted that the present invention is not limited to a real-time wide-angle video communication system, and the real-time wide-angle communication system described herein is only for illustrative purposes because the real-time wide-angle video communication system is typical of network bandwidth consumption. In fact, the present invention can also be implemented with any video communication system.

[0026] Figure 1 FIG. is an example diagram of a media transmission system 100 (including real-time wide-angle video transmission) drawn according to an embodiment of the present invention. As Figure 1As shown, system 100 may include multiple devices and networks, such as device 102, device 104, and network 106. These devices can be any configuration of one or more computers, such as a microcomputer, mainframe computer, supercomputer, general-purpose computer, special-purpose or dedicated computer, integrated computer, database computer, remote server computer, personal computer, laptop, tablet, mobile phone, personal digital assistant (PDA), wearable computing device, etc., or they can also be computing services provided by a computing service provider (such as web hosting or cloud services). In some embodiments, the computing device can be implemented by a combination of multiple sets of computers, and each computer can be located in a different geographical location and communicate with each other through a network or the like. Although certain operations can be jointly completed by multiple computers, in some embodiments, different computers will be assigned different operations. In some embodiments, system 100 can be implemented using a general-purpose computer or processor with a computer program, and when the computer program is running, it can execute the corresponding methods, algorithms, and / or instructions described in the present invention. Additionally, a special-purpose computer or processor equipped with special hardware can also be used to execute any method, algorithm, or instruction described in the present invention.

[0027] Device 102 may include internal hardware configurations such as processor 108 and memory 110. Processor 108 can be any type of one or more devices capable of operating or processing information. In some embodiments, processor 108 may include a central processing unit (such as a central processing unit, i.e., CPU). In other embodiments, processor 108 may include a graphics processing unit (such as a graphics processing unit, i.e., GPU). Although the examples described in the present invention can be implemented with the single processor shown, using multiple processors will be more advantageous in terms of speed and efficiency improvement. For example, processor 108 can be distributed across multiple machines or devices (each machine or device having one or more processors), and these machines or devices can be directly adapted and connected or interconnected through a network (such as a local area network).

[0028] Memory 110 can be any one or more temporary or non - temporary devices capable of storing code (e.g., instructions) and data, which can be accessed by a processor (via, for example, a bus). The memory 110 described in the present invention can be any combination of random access memory devices (RAM), read - only memory devices (ROM), optical or magnetic disks, hard disk drives, solid - state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, or any suitable type of storage device. In some embodiments, the memory 110 can be distributed across multiple machines or devices, such as network - based memory or cloud - based memory. The memory 110 can contain data (not shown in the figure), an operating system (not shown), and application programs (not shown). The data can be any data for processing (such as an audio stream, a video stream, or a multimedia stream). At least one application program can include a program that allows the processor 108 to execute instructions to generate control signals, which can be used to perform the various functions described in the methods below. For example, when the application program is acting as a sender, it can include computer instructions related to running Figure 5 the techniques described above.

[0029] In some embodiments, in addition to being equipped with the processor 108 and the memory 110, the device 102 can also include an auxiliary (e.g., external) storage device (not shown). If the above - mentioned auxiliary storage device is used, it can provide additional storage space during high - processing - demand situations. The auxiliary storage device can be a storage device in the form of any suitable non - temporary computer - readable medium, such as a memory card, a hard disk drive, a solid - state drive, a flash drive, or an optical drive, etc. Additionally, the auxiliary storage device can be either a component of the device 102 or a shared device accessed via a network. In some embodiments, the application programs in the memory 110 can be stored in whole or in part in the auxiliary storage device and loaded into the memory 110 as needed for processing.

[0030] In addition to including a processor 108 and a memory 110, the device 102 may further include input / output (I / O) devices. For example, the device 102 may include an I / O device 112. The I / O device 112 can be implemented in various ways. For instance, it can be a display adapted to the device 102 and configured to display images of graphic data. The I / O device 112 can be any device that transmits visual, auditory, or tactile signals to the user, such as a display, a touch-sensitive device (e.g., a touch screen), a speaker, headphones, a light-emitting diode (LED) indicator, or a vibration motor. The I / O device 112 can also be any type of input device that requires or does not require user intervention, such as a keyboard, a numeric keypad, a mouse, a trackball, a microphone, a touch-sensitive device (such as a touch screen), a sensor, or a gesture-sensing input device. If the I / O device 112 is a display, it can be a liquid crystal display (LCD), a cathode ray tube (CRT), or any other output device capable of providing a visible output to an individual. In some cases, an output device can also serve as an input device, such as a touch screen display that receives touch-based input.

[0031] The I / O device 112 can also consist of a communication device that transmits signals and / or data. For example, the I / O device 112 can include a wired device that sends signals or data from the device 102 to another device. Additionally, the I / O device 112 can also include a wireless transmitter or receiver using a compatible protocol for sending signals from the device 102 to another device or receiving signals from another device to the device 102.

[0032] The device 102 may further include a communication device 114 for communicating with another device. The communication between devices can be achieved through a connection to a network 106. The network 106 can be any suitable type of one communication network or any combination of multiple communication networks, including but not limited to Bluetooth communication, infrared communication, near field connection (NFC), wireless networks, wired networks, local area networks (LAN), wide area networks (WAN), virtual private networks (VPN), cellular data networks, and the Internet. The communication device 114 can be implemented in various ways, such as a transceiver device, a modem, a router, a gateway, a circuit, a chip, a wired network adapter, a wireless network adapter, a Bluetooth adapter, an infrared adapter, an NFC adapter, a cellular network chip, etc., or any combination of any suitable type of devices adapted to the device 102 to provide the function of communicating with the network 106.

[0033] The device 104 is similar to the device 102 and may include a processor 116, a memory 118, an I / O device 120, and a communication device 122. The components 116 - 122 provided in the device 104 can be similar to the corresponding components 108 - 114 in the device 102.

[0034] During different time periods of a real-time communication session, device 102 and device 104 can respectively act as a receiving device (i.e., the receiving end) or a transmitting device (i.e., the transmitting end). The receiving end can perform decoding operations, such as decoding a wide-angle video stream. Therefore, the receiving end can also be referred to as a decoding device or a decoding apparatus, and can be a decoder or a device containing a decoder. The transmitting end can also be referred to as an encoding device or an encoding apparatus, and can be an encoder or a device containing an encoder. Device 102 can communicate with device 104 via network 106.

[0035] Figure 2A It is a schematic diagram of image slices of a wide-angle video stream (i.e., video stream 200) in an embodiment. Video stream 200 can be a video source stream for encoding or a video stream decoded from a video bitstream. Video stream 200 shows two images (i.e., image 202 and image 204) of video stream 200. Image 202 and image 204 can be part of a larger overall image (such as a wide-angle image). That is, each of images 202 and 204 is divided into 10 slices, including slices 206 - 216, arranged in a 2x5 grid, and these slices can represent a viewport (i.e., a part) of the entire wide-angle image.

[0036] For example, slices 206 - 216 can be arranged in a 2x5 grid, and the entire wide-angle image can also be divided into a 5x5 grid. Therefore, the viewport represented by slices 206 - 216 can only represent 10 out of 25 slices of the entire wide-angle image. Thus, bitrate allocation can be applied to all slices of the entire wide-angle image, or only to a part of the slices, such as slices 206 - 216 representing the viewport of the overall wide-angle image.

[0037] It can be understood that video stream 200 can also contain more than two images. Additionally, as described above, images 202 and 204 are divided into slices. Each slice corresponds to a spatial position within the image, and the slices can be identified based on this position. Video stream 200 shows that each of images 202 and 204 is divided into 10 slices, segmented into a 2×5 grid, including slices 206 - 216. However, the present invention is not limited to this, and the image can be divided into more or fewer slices, rows, and / or columns.

[0038] In some examples, Cartesian coordinates can be used to identify slices. For example, the coordinate position of slice 206 can be considered as (0, 1), the coordinate position of slice 216 can be considered as (1, 3), and so on. In another example, identification can be performed according to the position of the slice in the image scanning order. That is to say, the slices can be numbered, for example, from 0 to the maximum number of slices in the image minus 1. Therefore, the slices of image 202 can be numbered from 0 to 9, where slices 206 - 216 are numbered as slices 1, 2, 3, 6, 7, and 8 respectively.

[0039] In video stream 200, in chronological order, image 204 can be after image 202. Video stream 200 shows a viewport 218, which contains an object 220 that a viewer (not shown) may be tracking. Viewport 218 is located within the enclosure of slices 206, 208, 212, and 214. Image 204 shows that object 220 has moved and the new viewport is now viewport 218', and viewport 218' is only located within slice 216.

[0040] Rate allocation can be used to encode a wide - angle video stream. Each image of the wide - angle video can be divided into slices. The rate can be allocated to the slices based on a total rate quota, and the total rate quota can be determined based on the total rate limit for the sender to transmit the image. That is to say, the total rate quota can be determined by the network capacity of the network. In this way, each slice can be allocated a part of the total rate quota. The allocated rate of each slice can be adjusted before, during, or after image transmission to improve the image quality of the image displayed by the receiver, as described below.

[0041] For example, as described above, image 204 can be divided into a 2×5 grid, containing a total of 10 slices, namely slices 206 - 216. Each slice can be allocated an initial rate for transmitting the slice from the sender to the receiver. For example, an equal total rate quota (e.g., 1 / 10 of the total rate quota) can be initially allocated to each slice. Image 204 can be sent from the sender to the receiver, and the sender and / or the receiver can evaluate the image quality of image 204. According to the evaluation, the sender and / or the receiver may consider that the image quality of image 204 needs to be improved. That is to say, the image 204 displayed by the receiver may be blurry, distorted, or unclear for the viewer.

[0042] To improve the image quality of image 204 and / or the images (i.e., subsequent images) transmitted from the sending end to the receiving end at a time point after the transmission of image 204, the receiving end can reallocate the total bitrate quota to slices. For example, since object 220 is located within slice 216, the total bitrate quota can be reallocated to increase the bitrate allocated to slice 216 while decreasing the bitrate allocated to one or more other slices. That is to say, due to the complexity of the image contained within slice 216 (e.g., because object 220 is located within slice 216), slice 216 may require a higher bitrate compared to other slices (e.g., those not containing object 220). Therefore, the overall bitrate quota can be dynamically allocated during a wide-angle video stream to improve the image quality displayed at the receiving end.

[0043] Figure 2B is a schematic diagram of slice encoding in an image of a wide-angle video stream. In Figure 2B there is a timeline, and the arrow indicates the direction of time. The slice stream 250 can include a series of slices along the timeline, including slices 252 - 258. The slice stream 250 can be Figure 2A the slice stream of slice 216 changing over time in

[0044] Each slice of the slice stream 250 can be divided into multiple processing units. In some video coding standards, the processing unit can be referred to as a "macroblock" or "coding tree block" (CTB). In some embodiments, each processing unit can further be divided into one or more processing subunits, where the processing subunits can be referred to as "prediction blocks" or "coding units" (CUs) according to different criteria. The size and shape of the processing unit and subunits can be arbitrary, such as 8×8, 8×16, 16×16, 32×32, 64×64, etc., or any shape and size suitable for encoding an image region. Generally, the more details an encoding region contains, the smaller the size of the processing unit and subunits can be divided. For ease of understanding and without causing ambiguity, the processing unit and subunits are collectively referred to as "tiles" hereinafter, unless otherwise specifically stated. For example, in Figure 2B slice 256 is divided into 4×4 tiles, including tile 210. The boundaries of the tiles are indicated by dashed lines. Of course, the slicing of the slices in the slice stream 250 is not limited to the above example, and each slice can be subdivided in any desired manner according to the description herein.

[0045] Figure 3 is a schematic diagram of a network 300 for real-time video communication in an embodiment, where the real-time video communication includes wide-angle video communication. The network 300 can be in a computing network (such asFigure 1 implemented on the application layer of the network 106). For example, in the TCP / IP model, a computer communication network can be divided into multiple layers. Again, in the hierarchical order from bottom to top, it can include a physical layer, a network layer, a transport layer, and an application layer respectively. Each of the above layers provides services for the layer above it and is provided with services by the layer below it. The application layer is the TCP / IP layer that directly interacts with the end user through software applications. In Figure 1 the network 106, the network 300 can be implemented and deployed in the form of an application layer software module.

[0046] In some embodiments, the network 300 can be software installed on a node (such as a server) of the network 106. In some embodiments, the network 300 does not require the installation of dedicated or specific hardware (such as dedicated or proprietary network access point hardware, etc.) on the node. For example, the node can be any x86 or x64 computer installed with an operating system (OS), and the network interface at the node serving as the access point of the network 300 can be any general network interface hardware (such as an RJ-45 Ethernet adapter, a wired or wireless router, a Wi-Fi communication adapter, or any general network interface hardware, etc.).

[0047] In addition, the network formed by the network 300, that is, Figure 1 the network 106 in

[0048] can also be a public network (such as the Internet). In some embodiments, the nodes of the network 300 can communicate with each other through the network 300 or the public network. In other words, part of the data traffic of the network 300 can be routed via the public network instead of all within the network 300. In some embodiments, all the nodes in the network 300 can simultaneously communicate with data traffic interaction through the network 300 and the public network.

[0048] As Figure 3 shown, the network 300 includes two types of nodes: service nodes and control nodes. Service nodes (such as service nodes 304 - 318) are used to receive multimedia data from different user terminals, cache multimedia data, forward multimedia data, and transmit multimedia data to different user terminals. Service nodes can also receive and send feedback messages, as detailed below. Control nodes (such as control node 302) can be used to control network traffic. Service nodes and control nodes can be interconnected, although it is not fully shown in Figure 3 That is to say, any two nodes in the network 300 can be directly connected. The connection between nodes can be bidirectional or unidirectional. The connection between nodes can be partially bidirectional and partially unidirectional. As described above, there may not be a direct connection between two nodes, so two nodes can also communicate indirectly through a third node.

[0049] AsFigure 3 As shown, service nodes can be subdivided into two types: edge service nodes (abbreviated as "edge nodes" or "edge devices") and router service nodes (abbreviated as "routing nodes"). Edge nodes are directly connected to end-user terminals (abbreviated as "terminals"), such as terminals 320-326. Terminals include any end-user device capable of multimedia communication, such as smartphones, tablets, cameras, monitors, laptops, desktop computers, workstation computers, or devices with multimedia I / O. Routing nodes may not be directly connected to any terminal. Routing nodes (such as service nodes 304, 308, and 312-316) participate in forwarding data. In some embodiments, a service node can switch roles between an edge node and a routing node at different times, or act as both simultaneously.

[0050] It should be noted that network 300 can have any number of any type of nodes and have any interconnection configuration, not limited to Figure 3 as shown. It should also be noted that in the examples herein, network 300 includes service nodes 304-318 and control node 302, but according to the description of the present invention, any other network configuration can also be adopted. That is to say, network 300 is only an illustrative example, and the method described in the present invention can also be implemented by any type of network (for example, the Figure 4A-5 bitrate allocation method described below).

[0051] For example, the present invention can be implemented by any type of network, including but not limited to: point-to-point communication networks, multi-point communication networks, broadcast networks, video conferencing systems, satellite communication networks, Internet Protocol Television (IPTV), mobile video communication or point-to-point video communication using cellular networks and / or Wi-Fi, local area networks (LANs), and wide area networks (WANs). Therefore, network 300 can be considered as a certain type of network used for describing the present invention.

[0052] Figure 4A is a flowchart of a technique 400 for allocating bitrates for video communication (such as real-time wide-angle video communication) in an embodiment. Technique 400 can be implemented by a sending end (such as Figure 1 devices 102 and / or 104). Technique 400 can be implemented as a software module stored in Figure 1 memory 110 and / or memory 118 and respectively as instructions and / or data run by processors 108 and 116 in Figure 1 . Again, for example, technique 400 can also be implemented as hardware, such as a dedicated chip that can run dedicated instructions.

[0053] Technique 400 can be executed by the sender at each time step. For example, if the sender is transmitting images of a wide-angle video for display at a rate of 30 frames per second, Technique 400 can be executed approximately every 33 milliseconds. And so on, Technique 400 can be executed by the sender according to a predefined duration. For example, Technique 400 can be executed by the sender every X time steps, where X can be defined as a set number of time steps. Additionally, it should be noted that all or part of Technique 400 can also be executed by the receiver. For example, a part of Technique 400 can be executed by the receiver and transmitted to the sender via a network, such as Network 106.

[0054] As described above, the sender can be configured to encode the images of the video stream using slices, whereby each slice can be allocated a bitrate based on the total bitrate quota. Allocating bitrates to slices can be done dynamically. That is, the total bitrate quota can initially be evenly allocated to each slice, and after evaluating one or more factors, the sender can reallocate the total bitrate quota to the slices. In this way, the initial bitrate allocated to a slice can be modified to increase and / or decrease the bitrate, thereby improving the image quality displayed by the receiver.

[0055] As described above, bitrate allocation can be done within each time step of the video stream, or it can also be done according to a predefined time interval. The time interval for bitrate allocation can be predefined or can be dynamically adjusted based on the evaluation of the video stream by the sender and / or the receiver, as detailed below. For example, when the image quality of the image displayed by the receiver is poor, the frequency of bitrate allocation can be increased, while when the image quality of the image displayed by the receiver is clear, bitrate allocation can be done less frequently.

[0056] Technique 400 can include one or more operation modes. The operation modes of Technique 400 can include one or more steps (i.e., operations) of Technique 400. The steps included in an operation mode can be similar or can be unique to a particular operation mode. For example, Technique 400 can include two operation modes, and the two operation modes can include the same steps or shared steps. However, each operation mode can include unique steps that the receiver does not execute in both operation modes.

[0057] As Figure 4A shown, Technique 400 can include a bitrate iteration mode 402 and an image quality monitoring mode 420. Technique 400 can also switch between the bitrate iteration mode 402 and the image quality monitoring mode 420, as detailed below.

[0058] The bitrate iteration mode 402 starts from step 404. That is, the startup technique 400 starts with step 404 in the bitrate iteration mode 402. After the bitrate iteration mode 402 starts at step 404, the video stream can be encoded at step 406 for transmission. As described above, the technique 400 can be used for bitrate allocation of a video stream (such as a real-time wide-angle video stream). The video stream can be divided into slices associated with the viewing angle of the video stream. Based on the viewing angle of the video stream, the slices can be associated with the images of the video stream at each time step. That is, the video stream can be encoded and sent from the sender to the receiver as images divided into slices over a period of time.

[0059] The encoding at step 406 can include encoding each slice (such as Figure 2A slice 206 - 216 in). The encoding at step 406 can include allocating a bitrate for each slice. The encoder (i.e., the sender) can allocate a portion of the total bitrate quota (i.e., the total bitrate amount from the sender to the receiver) according to the bitrate allocation bitrate vector. The bitrate allocation bitrate vector (b j ) is in units of kilobits per second (kbps) and is obtained using the following formula:

[0060] b j = Bx j-1 (Formula 1)

[0061] where B is the total bitrate quota (kbps), x j-1 is the bitrate allocation weight vector, and j is the iteration round. The iteration of the technique 400 within the bitrate iteration mode 402 can be considered as the process of completing the loop of steps 406 - 416. For example, each iteration can include encoding the slices at step 406 and evaluating one or more of the images sent to the receiver, and then subsequent iterations can return to step 406 to start reallocating the bitrate to the slices.

[0062] The initial state (i.e., the initial iteration with j = 0) can be determined when the user switches to a new viewing angle of the video stream. In this initial state, all slices associated with the new viewing angle can be encoded with an initial bitrate allocation. Assuming the total number of slices transmitted from the sender to the receiver is n, the initial bitrate allocation weight is calculated as follows:

[0063]

[0064] where the n slices can be allocated equal bitrates based on the total bitrate quota B. That is, the bitrate of each slice can be

[0065] After encoding is completed at step 406, the sender can send the encoded slices to the receiver. Then the receiver can slice render at the first iteration (i.e., the initial iteration where j = 0). The receiver's rendering can include decoding and / or displaying the slices. For example, the receiver can receive the encoded slices with an initial bitrate allocation, decode the slices, remap and stitch the slices, and then display the slices to present a view of the video (e.g., display the first image of the video stream). It should be noted that the receiver can also receive and decode the slices but not display them. For example, only the slices within the viewer's viewport are sent to the receiver. In this case, all slices of the image are available for rendering and display, however, since parts of some slices are outside the viewport, these slices may only be partially rendered.

[0066] After the slices are encoded at step 406, the encoded slices are sent to the receiver and an image quality score can be determined at step 408. Determining the image quality score at step 408 can include evaluating the quality of the image sent to the receiver (i.e., the image divided into encoded slices). That is, the image quality score can be a comparison of the compressed (i.e., encoded) image and the original image. The image quality score can be determined by the sender and / or the receiver. For example, the image quality score can be determined by the sender by comparing the quality of the encoded (i.e., compressed) image or its slices with the quality of the original, uncompressed image or its slices. Or, the image quality score can be directly determined by the receiver by evaluating the rendered (i.e., decoded and / or displayed) image or slices using a no-reference image quality assessment model. Additionally, the sender and the receiver can communicate with each other to determine the image quality score. For example, the receiver decodes and displays the image, at which time the receiver can send information related to the displayed image back to the sender, and then the sender uses the information sent from the receiver to evaluate the image quality score of the image.

[0067] The image quality score can be determined using various metrics and / or subjective evaluations. That is, determining the image quality score is not specifically limited to any one method. Objective metrics can be used to determine the image quality score, such as by mean squared error (MSE), peak signal-to-noise ratio (PSNR), or structural similarity index (SSI), etc. The image quality score can also be determined by evaluating various metrics of the image, such as contrast, brightness, and color accuracy, etc. Additionally, one or more machine learning models (e.g., deep learning models of neural networks) can be used to determine the image quality score.

[0068] The image quality score of the encoded image can be calculated at each time step of the video stream or can be calculated according to a predefined time interval. That is, the image quality score can be determined according to a defined time interval, that is, the image quality score can be determined for each time step or a part of the time step. At step 408, the overall image quality score can be calculated based on the previously determined image quality scores. For example, the overall image quality score can be calculated according to the image quality scores calculated for the previous iterations of the video stream in the past T seconds. That is, the overall image quality score can be calculated according to the image quality scores calculated for the previous iterations within a set time range (i.e., the past T seconds). Therefore, the overall image quality score (q j ) can be expressed as:

[0069]

[0070] where j represents the iteration round of technology 400 within the bitrate iteration mode 402, and n represents the number of slices. For example, as Figure 2A shown, n can be a number from 1 to 10, representing 10 slices of image 202, including slices 206 - 216.

[0071] After determining the overall image quality score (q j ), the average image quality score of all slices of the image can be determined using the following formula

[0072]

[0073] After determining the average image quality score using Equation 4 , the image quality gain can be calculated at step 410. The image quality gain represents the average gain (e.g., improvement) in the image quality of the encoded image from the sender to the receiver. For example, the image quality gain can represent the average gain (e.g., improvement) in the image quality score, which is determined according to the set change in the bitrate between iterations j in the above example. This set change in the bitrate can be caused by the change in the bitrate allocation of the slices. For example, the image quality gain can represent the average gain in the image quality score caused by a 100 kbps change in the bitrate between iterations j. Therefore, the image quality gain can represent the improvement rate of the image quality, which is caused by the increase in the slice bitrate due to the reallocation of the bitrate for each slice. That is, the image quality gain can track the change rate (e.g., improvement rate) of the image quality between iterations j of the bitrate iteration mode 402.

[0074] Before calculating the image quality gain, it can be first determined whether the ratio of the bitrate allocation between slices has changed between the current iteration and the previous iteration. That is, the change in the bitrate allocation can be represented by the following formula:

[0075] x j != x j-1 (Formula 5)

[0076] Where, as described above, x represents the rate allocation weight vector for allocating the bitrate to the slices.

[0077] If the bitrate allocation ratio changes as described above, the average image quality gain gain j can be obtained by the following formula:

[0078]

[0079] Where m is the number of slices that get an increase in bitrate, is the bitrate (kbps) of the i-th slice in the j-th iteration. As described above, if the bitrate allocation ratio does not change (e.g., x j = x j-1 ), then the gain can be determined to be 0, or it means that the transmitter and / or receiver provide invalid results.

[0080] After determining the image quality gain at step 410, it can be determined at step 412 whether to stop the iteration of the bitrate iteration mode 402. That is, it can be determined at step 412 whether the technique 400 can switch from the bitrate iteration mode 402 to the image quality monitoring mode 420. Since the bitrate is dynamically reallocated to the slices in the bitrate iteration mode 402, the operation of the technique 400 in the bitrate iteration mode 402 may require a large amount of computing power and / or network bandwidth.

[0081] In contrast, the image quality monitoring mode 420 can only monitor the image quality without dynamically allocating the bitrate to the slices. Moreover, compared with the bitrate iteration mode, the computing frequency can be lower. In this way, compared with the bitrate iteration mode 402, the image quality monitoring mode 420 may require less computing power and / or network bandwidth. For example, as described below, the image quality monitoring mode 420 can monitor the image quality. If the image quality is not within the acceptable range or does not meet one or more parameters, the technique 400 can switch back from the image quality monitoring mode 420 to the bitrate iteration mode 402, and at this time, the bitrate can be evaluated and reallocated according to the above method.

[0082] As Figure 4BAs shown, one or more factors can be evaluated at step 412 to determine whether the bitrate iteration mode 404 should be stopped and whether the technique 400 should switch to the image quality monitoring mode 420. For example, if the image quality gain is too low, the technique 400 can switch from the bitrate iteration mode 404 to the image quality monitoring mode 420. Additionally, if the image quality converges between slices such that the maximum image quality slice and the minimum image quality slice fall within a threshold, the technique 400 can switch from the bitrate iteration mode 404 to the image quality monitoring mode 420. Further, if the bitrate of any slice reaches a preset minimum bitrate, the technique 400 can switch from the bitrate iteration mode 404 to the image quality monitoring mode 420. However, any factor can be evaluated at step 412 to determine whether to switch from the bitrate iteration mode 402 to the image quality monitoring mode 420.

[0083] If at step 412 it is determined based on the above factors that the technique 400 should switch from the bitrate iteration mode 402 to the image quality monitoring mode 420, an image quality score can be calculated at step 424. The method for calculating the image quality score can be similar to the method for calculating the image quality score at step 408. It should be noted that during the image quality monitoring mode 420, the bitrate allocation ratio between slices (i.e., the bitrate allocated to each slice) can remain the same. The calculation interval between each image quality score calculation can also be increased. That is, the time interval (i.e., duration) between the image quality score calculations at step 424 can be greater than the time interval (i.e., duration) between the image quality score calculations at step 404 in the bitrate iteration mode 402. The duration can be any predefined or adjusted time interval according to the technique 400.

[0084] After determining the image quality score at step 424, the image quality change can be determined at step 426. That is, the image quality score determined at step 424 can be compared with the image quality score from the previous or most recent iteration j. Depending on the current iteration of the technique 400, the previous or most recent iteration j can be calculated at step 408 or 424.

[0085] After determining the image quality change, it can be determined at step 428 whether the technique 400 should switch from the image quality monitoring mode 420 to the bitrate iteration mode 402. The above determination can be based on the image quality change determined at step 426. For example, if the image quality score vector after the last iteration is q l (i.e., the image quality score vector determined at step 408 during the last iteration of the bitrate iteration mode 402), and the image quality score vector in the current image quality monitoring mode 420 is q j, the following formula can be used to determine whether to switch back to the bitrate iteration mode 402 at step 428:

[0086]

[0087] If it is determined that Equation 7 is true, that is, if the change in image quality determined according to the above method is greater than the threshold (q thres ), then the technique 400 can switch from the image quality monitoring mode 420 to the bitrate iteration mode 402. However, if the change in image quality determined according to the above method is less than the threshold (q thres ), then the technique 400 remains in the image quality monitoring mode 420 and recalculates the image quality score at step 424 after a predefined calculation interval.

[0088] If the sender and / or receiver determines not to enter the image quality monitoring mode 420 at step 412, or if the sender and / or receiver determines to re-enter the bitrate iteration mode 402 at step 428, then the iteration step vector can be determined at step 414. The iteration step vector determined at step 416 can be derived from the bitrate allocation weight vector x j , or used in combination with the bitrate allocation weight vector x j . Specifically, the iteration step vector determined at step 416 can be used to determine the new bitrate for each slice (i.e., for reallocating the new bitrate).

[0089] For example, as described above, the image quality of the encoded image (e.g., encoded slice) can be improved at each iteration of the bitrate iteration mode 402, thereby improving the user experience when the encoded image is sent from the sender to the receiver and displayed. This improvement in image quality (e.g., improvement in the image quality score) can be accomplished by reallocating the total bitrate quota to adjust the bitrate of each slice. In this case, the bitrate of low-complexity slices that contain the fewest or no objects (such as slices 206 - 214 of image 204) can be reduced. Since the bitrate of the low-complexity slices is reduced, the bitrate of high-complexity slices that contain more details and / or objects (such as slice 216 in image 204) can be increased. That is, the remaining bitrate that is not needed to encode and transmit the low-complexity slices with an acceptable image quality can be reallocated to the high-complexity slices that require more bitrate for encoding and transmission with an acceptable image quality.

[0090] It should be noted that although the above allocation method is generally acceptable for improving image quality, under certain conditions, the rate of increase and / or decrease of the slice bitrate may be too significant, thus having a negative impact on the user experience (e.g., having a negative impact on video quality). For example, the aforementioned low-complexity slices can generally obtain a higher image quality score at step 408. Due to their high image quality scores, technique 400 may significantly reduce their bitrate to increase the bitrate of high-complexity slices. The iteration step vector determined at step 414 can be used to ensure that the rate of decrease in this case is not too drastic, that is, to ensure that the gradient of decrease between the bitrate of the current iteration and the bitrate of subsequent iterations is not too large.

[0091] For example, the iteration step vector (w j ) can be used to reduce the variance of the image quality among all slices. The bitrate of each slice can be proportional to the image quality of each slice. According to such a proportion, the bitrate allocation weight (x j ) can be determined in the following manner:

[0092]

[0093] where q j represents the image quality score at the j-th iteration, represents the average image quality score of all image slices at the j-th iteration, and α represents the control parameter of the iteration step. It should be noted that the control parameter (α) can be a default value (e.g., 0.5) and can be adjusted according to the needs of technique 400.

[0094] In addition to determining the bitrate allocation weight (x j ) according to Equation 8, the changed slice bitrate step (e.g., the difference in bitrate allocated between iterations) can be proportional to the following:

[0095]

[0096] Based on the bitrate allocation weight and the slice bitrate step size determined according to the above formula, the step vector (w j ) can be determined. All elements in are greater than 0 and can be represented as a vector The average value of the vector can be represented as In addition, the vector can be readjusted to reduce the variance between the image quality scores of the slices, such that the vector is represented as a vector where the vector can be represented according to the following formula:

[0097]

[0098] where β is between 0 and 1, which can make the original vector converge to the average image quality of the slices.

[0099] In summary, the positive vector in can be replaced by to obtain the final iteration step size vector (w j ).

[0100] For example, the image can include a first slice, a second slice, and a third slice. The first vector value, the second vector value, and the third vector value can measure the differences between the average image quality of all slices and the first slice, the second slice, and the third slice, respectively. For example, the first vector value can measure the difference between the image quality of the first slice and the average image quality as 1.5; the second vector value can measure the difference between the image quality of the second slice and the average image quality as 0.8; the third vector value can measure the difference between the image quality of the third slice and the average image quality as 0.6. Assuming that the first slice, the second slice, and the third slice all have higher image quality than the average image quality (i.e., the first slice, the second slice, and the third slice are all positive values), the vector can be represented by the following formula:

[0101]

[0102] Then, the above vector values of the first slice, the second slice, and the third slice can be adjusted using Equation 9 above (i.e., using β in Equation 9, where β can have a value between 0 and 1) to reduce the variance between the vector values of the first slice, the second slice, and the third slice, thereby determining the readjusted vector For example, if β is 0.5, the vector can be readjusted according to Equation 9 as:

[0103]

[0104]

[0105] Thus, the readjusted vector replaces to obtain the final step size vector (w j ).

[0106] After determining the iteration step size vector at step 414, the new rate weight (e.g., rate allocation weight vector) for each slice can be determined at step 416. The new rate allocation weight vector (x j ) can be determined using the following formula:

[0107] ​

[0108] Where α is a parameter that controls the step size of iteration j. As described above, the larger the value of α, the more drastic the change in the bitrate of the slice, and vice versa.

[0109] As described above, the change in the bitrate of the slice in the current iteration compared to the previous iteration at step 416 can be gradually accomplished (e.g., in a gradient ascent or descent) without compromising the viewing experience of the user for the image displayed at the receiving end. To further ensure that the bitrate of the slice is dynamically allocated in a progressive manner, a validity check can be performed at step 416. By way of example, at step 416, the following formula can be used to determine whether the bitrate allocation weight calculated at step 416 is valid:

[0110] max(q j ) - min(q j ) ≤ quality_margin (Formula 13)

[0111] That is, a validity check can be performed at step 416 based on determining whether the difference between the maximum image quality score of the slice and the minimum image quality score of the slice is less than or equal to a defined threshold (i.e., quality_margin). If the validity check at step 416 is successful and the condition of Formula 13 is met, the bitrate allocation weight can be determined as x j = x j-1 . After the condition of Formula 13 is met, the bitrate allocation weight can be used to encode the slice in the next iteration at step 406. Conversely, if the condition of Equation 13 is not met, i.e., the difference between the maximum image quality score of the slice and the minimum image quality score of the slice is greater than the defined threshold (i.e., quality_margin), the bitrate allocation weight and the updated value obtained through Formula 12 can be returned together.

[0112] In addition, it should be noted that after any iteration of the bitrate iteration mode 402 is completed, the bitrate iteration mode can end at step 418. That is, the technique 400 can end at step 418 due to a user input or the termination of the video stream. For example, if a real-time wide-angle video stream is aborted at the sending end, the technique 400 can end at 418 or otherwise be completed.

[0113] Although the paths 404 - 418 and 424 - 428 shown in the technique 400 can be executed in an alternating manner, this is not necessarily the case. In some embodiments, the path 404 - 418 can be executed in parallel with the path 424 - 428. That is, in some embodiments, the image quality monitoring mode 420 can be executed in parallel with the bitrate iteration mode 402. In addition, the technique 400 can be deployed in other ways.

[0114] Figure 4B is a schematic diagram of an operating mode of a technique 400 for allocating bitrates for video communication, such as real-time wide-angle video communication. As described above, the technique 400 can be implemented as a software module stored in Figure 1 the memory 110 and / or the memory 118, and respectively as instructions and / or data executed by the Figure 1 processor 108 and the processor 116 in. For example, the technique 500 can be implemented in hardware as a dedicated chip storing instructions executable by the dedicated chip. Additionally, as described above with respect to the technique 400, the technique 500 can be executed by the sending end and / or the receiving end at each time step.

[0115] As described above, the technique 400 can include a bitrate iteration mode 402 and an image quality monitoring mode 420. The technique 400 can determine whether to switch between the bitrate iteration mode 402 and the image quality monitoring mode 420 based on whether one or more conditions are met. Determining whether one or more conditions are met can be done at step 412 or step 428. That is, determining whether one or more conditions are met can be done at step 412 in the bitrate iteration mode 402 to stop the iteration and enter the image quality monitoring mode 420; or determining whether one or more conditions are met can also be done at step 428 in the image quality monitoring mode 420 to restart (i.e., re-enter) the bitrate iteration mode 402.

[0116] Switching from the bitrate iteration mode 402 to the image quality monitoring mode 420 helps reduce the frequency of evaluating image quality, thereby reducing the required computing power and / or network bandwidth of the network. To determine whether to enter the image quality monitoring mode 420 and stop the iteration of the bitrate iteration mode 420 at step 412, one or more conditions can be evaluated. If all or some of the conditions are met, the iteration can be stopped at 412.

[0117] For example, if the image quality gain threshold is not reached at step 430, the iteration can be stopped at step 412 and the technique 400 can enter the image quality monitoring mode 420. At step 430, it can be determined whether the average image quality gain of a picture (e.g., a slice) is too low compared to a predefined threshold. If the average image quality gain is too low, the image quality monitoring mode 420 can be entered according to the path 424 - 426.

[0118] To evaluate this condition at step 430, it can first be determined whether the bitrate has been reallocated within the current iteration. If the bitrate has been reallocated within the current iteration j, i.e., x j != x j-1, the average gain (gain j ) of the image quality score of an image (e.g., a slice) can be obtained by the following formula:

[0119]

[0120] where m represents the number of slices for which the bitrate has increased (i.e., the number of slices for which the reallocated bitrate is greater than its bitrate in the previous iteration), j represents the j-th iteration, and i represents the i-th slice among all m slices. In summary, the average gain (gain j ) can be the average gain of the image quality score per 100,000 bitrate changes after this iteration.

[0121] As described above, the following method can be used to compare the average gain with the gain threshold:

[0122] gain j > gain_thresh (Formula 15)

[0123] If the average gain is greater than the threshold, then Technique 400 may not stop the iteration at step 412 and may continue the bitrate iteration mode 402 in the next iteration along path 406 - 416. However, if the average gain is less than the threshold, then the iteration can be stopped at step 412 and enter the image quality monitoring mode 420. It should also be noted that the above gain threshold evaluation may require a single calculation result of Formula 15, or may also require consecutive multiple calculation results of Formula 15 to determine whether to stop the iteration at step 412. For example, the gain threshold can be evaluated at step 430 in each iteration, and the average gain that is continuously less than the threshold must be returned in a predetermined number of iterations (e.g., 2 or 3 iterations) before entering the image quality monitoring mode 420.

[0124] In addition to the image quality gain threshold evaluation at step 430, a bitrate threshold can also be evaluated at step 432 to determine whether to stop the current iteration at step 412. The bitrate evaluation at step 430 can be obtained by comparing the minimum allocated bitrate b j with the threshold:

[0125] min(b j ) ≤ bitrate_thresh (Formula 16)

[0126] If the minimum allocated bitrate is less than or equal to the threshold (bitrate_thresh), then the iteration can be stopped at step 412 and enter the image quality monitoring mode 420. Conversely, if the minimum bitrate allocation bitrate is greater than the threshold (bitrate_thresh), then it can continue in the bitrate iteration mode 402.

[0127] Additionally, it can be determined at step 434 whether the image quality of the slices converges. That is, if the difference between the maximum and minimum values of all slice image quality scores (q max and q min ) is equal to or falls within a defined threshold, the iteration can stop at step 412. Similar to the image quality gain threshold evaluation at step 430, the image quality convergence at step 434 can be evaluated in each iteration of the bitrate iteration mode 402, and the iteration may not stop at step 412 until the defined number of iterations meets the conditions for the image quality convergence evaluation. The image quality convergence can be evaluated according to the following formula:

[0128] q max -q min ≤bitrate_thresh (Formula 17)

[0129] Therefore, as described above, if the image quality gain threshold is not reached at step 430, if the bitrate threshold is not reached at 423, or if the image quality converges within the defined threshold at 434, the iteration mode 402 of the bitrate iteration can stop at step 412, and the image quality monitoring mode 420 can be entered.

[0130] As described above, the technique 400 can re-enter or restart the bitrate iteration mode 402 at step 428 based on one or more conditions. For example, it can be determined at step 436 whether to re-initiate the iteration of the bitrate iteration mode 402 at step 428 by comparing the image quality scores of the current iteration relative to the previous iteration. Specifically, the following formula can be used to determine the image quality change at step 436:

[0131]

[0132] where q j is the image quality score of the current iteration j, and q l is the image quality score of the previous iteration. If the quality threshold (q thresh ) is exceeded, the bitrate iteration mode 402 can be re-entered at 428, and the step of calculating the iteration step vector at 414 can continue to run within the iteration, as described above. Conversely, if the quality threshold (q thresh ) is not reached, the image quality monitoring mode 420 can continue to run along path 424 - 426, as described above.

[0133] Figure 5 is a flowchart of a technique 500 for allocating bitrates for video communication (such as real-time wide-angle video communication). The technique 500 is based on andFigure 4A and 4B is similar to the technique 400. The technique 500 can be implemented by a sending end (i.e., an encoder) (such as Figure 1 the device 102 and / or the device 104). The technique 500 can be implemented as a software module, which is stored in Figure 1 the memory 110 and / or the memory 118, and respectively as instructions and / or data executed by Figure 1 the processor 108 and the processor 116. For example, the technique 500 can also be implemented in hardware as a dedicated chip that stores instructions executable by a dedicated chip. Alternatively, a part of the technique 500 can also be implemented by Figure 3 a service node in the network 300. For example, a part of the technique 500 can be implemented as a software module stored in the memory of a network node, as instructions and / or data executable by a program of the network node or a program of the sending end (such as an encoder).

[0134] At step 502, the video stream can be encoded, whereby the video stream can be divided into slices, such as Figure 2A the slices 206 - 216. As Figure 2A shown, the slices can be associated with the viewpoints of the video stream, such as Figure 2A the image 204. Each slice can be assigned a certain bitrate at step 502_1. The bitrate assignment can refer to the encoding process described in step 406 of the technique 400 as described in Figure 4A . Similarly, the assignment can also be completed according to the re - assignment process shown in the path 406 - 416 in the technique 400. In either case, the encoding at step 502 can be completed within the iteration of the bitrate iteration mode 402 of the technique 400. In short, the encoding at step 502 can be completed at the first iteration of the bitrate iteration mode 402.

[0135] After encoding the slices at step 502, an image quality score can be determined at step 504. The image quality score can be associated with the rendering of the slices (e.g., slice rendering by the receiving end), so the rendering of the slices can include decoding and / or displaying the slices at the bitrate from the viewpoint of the video stream at the first iteration. The image quality score can be determined before or after the receiving - end slice rendering. For example, the image quality score can be determined by the sending end after encoding the slices, or the image quality score can also be determined by the receiving end after decoding and / or displaying the image. In addition, the image quality score can be determined with reference to step 408 of the technique 400 as described above.

[0136] After determining the image quality score at step 504, the bitrate allocated to one or more slices can be adjusted at step 506. Adjusting the bitrate allocated to one or more slices can include reallocating the bitrate of low-complexity slices to high-complexity slices, thereby improving the overall image quality presented at the receiving end. This reallocation can be accomplished by Technique 400.

[0137] For example, the process of determining the image quality score at step 504 can be similar to step 408 of Technique 400. After determining the image quality score at step 504, the bitrate of the slices at step 504 can be adjusted (e.g., reallocated). Based on the calculated image quality gain, the calculated iteration step vector, and / or the calculated bitrate allocation weight, step 506 adjusts the bitrate. The image quality gain, the iteration step vector, and the bitrate allocation weight can be calculated by paths 410-416 of Technique 400 as described above. Thus, the calculation results can be used to determine the bitrate reallocation of the slices.

[0138] After reallocating the bitrate for the slices, the second iteration (i.e., the iteration after the first iteration) can be encoded and sent by the transmitting end to the receiving end for rendering, such that the rendering of the slices in the second iteration includes decoding and displaying the slices with the adjusted (e.g., reallocated) bitrate. Therefore, compared with the image displayed at the first iteration, the image decoded and displayed at the second iteration can be improved. Technique 500 can be executed any number of times to improve the overall image quality by reallocating the bitrate to the slices. Additionally, Technique 500 can include all or part of the bitrate iteration mode 404 as described above and / or can include all or part of the image quality monitoring mode 420 as described above.

[0139] As described above, those skilled in the art should understand that all or part of the content described herein can be implemented using a general-purpose computer or processor with a computer program that, when run, can execute any corresponding techniques, algorithms, and / or instructions described herein.

[0140] The computing device described in the present invention, as well as the algorithms, methods, instructions, etc. stored thereon and / or executed thereby, can be implemented by hardware, software, or any combination thereof. The hardware can include a computer, an intellectual property (IP) core, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), an optical processor, a programmable logic controller, microcode, a microcontroller, a server, a microprocessor, a digital signal processor, or any other suitable circuit. In the claims, the term "processor" should be understood to encompass any of the above and can be one or a combination of several of them.

[0141] The various functions described in the present invention can be described by functional components and various processing operations. The processes and sequences described in the present invention can be executed individually or in any combination. The functional modules can be implemented by any number of hardware and / or software components that can run specific functions. For example, the content described above can adopt various integrated circuit components, such as memory elements, processing elements, logic elements, lookup tables, etc., which can execute various functions under the control of one or more microprocessors or other control devices. Similarly, for the content of the present invention implemented by software programming or software programs, any programming or scripting language such as C, C++, Java, assembly programs, etc. can be used to implement it, and any combination of any data structures, objects, processes, routines, or other programming elements can be used to execute various algorithms. The various functions can be implemented by executing algorithms on one or more processors. In addition, the various functions described in the present invention can adopt any number of conventional technologies for electronic configuration, signal processing and / or control, data processing, etc. The terms "mechanism" and "element" are widely used herein, but do not mean limited to being implemented by mechanical or physical methods, but can include software routines adapted to run on a processor, etc.

[0142] Embodiments or some embodiments of the present invention may take the form of a computer program product, which can be accessed through a computer-usable medium or a computer-readable medium, etc. The computer-usable medium or the computer-readable medium can be any suitable device, which can specifically contain, store, transmit, or transfer a program or data structure for use by or connected to any processor. The form of the medium can be electronic, magnetic, optical, electromagnetic, or semiconductor devices, etc., and of course can also include other applicable media. The above-mentioned computer-usable medium or computer-readable medium can be referred to as non-transitory memory or medium, and can include RAM or other volatile memory or storage devices, which can change over time. Unless otherwise specifically stated in the text, the device memory described herein is not necessarily physically equipped in the device, but can be remotely accessed by the device and does not have to be adjacent to other physically equipped memories in the device.

[0143] One or more functions executed in the embodiments of the present invention can be implemented by machine-readable instructions in the form of code, which are used to operate the above-mentioned one or more hardware combinations. The computing code can be implemented in the form of one or more modules, through which the functions of one or more combinations can be executed as computing tools. When the methods and systems described in the present invention are running, the input and output data are transmitted between each module and one or more other modules.

[0144] Terms such as "signal" and "data" are used interchangeably herein. In addition, the functions of the various parts of a computing device do not necessarily have to be implemented in the same way. Information, data, and signals can be represented using a variety of different technologies and methods. For example, any data, instructions, commands, information, signals, bits, symbols, and chips mentioned herein can be represented by one or a combination of voltage, current, electromagnetic waves, magnetic fields or particles, optical fields or particles, etc.

[0145] The present invention uses the term "example" to mean an illustration, instance, or exemplification. Any aspect or design described herein as an "example" is not necessarily meant to represent the best mode of the present invention. The term "example" is used to present concepts in a concrete manner. Additionally, the phrases "an aspect" or "one aspect" are used multiple times throughout the text, but do not necessarily mean the same embodiment or have the same function.

[0146] The word "or" as used in the present invention is intended to mean an inclusive "or" rather than an exclusive "or". That is, "X includes / contains A or B" is intended to mean any natural inclusive arrangement, unless otherwise stated or clearly determinable from the context. In other words, "X includes A or B" can mean any of the following: X includes A; X includes B; X includes A and B. By analogy, "X includes one of A and B" means "X includes A or B". The phrase "and / or" as used herein is intended to mean "and" or an inclusive "or". That is, "X includes A, B, and / or C" is intended to mean that X can include any combination of A, B, and C, unless otherwise stated or clearly indicated by the context. In other words, if X includes A, X includes B, X includes C, X includes A and B, X includes B and C, X includes A and C, or X includes all of A, B, and C, then each or any combination of the above cases satisfies the description of "X includes A, B, and / or C". By analogy, "X includes at least one of A, B, and C" is equivalent to "X includes A, B, and / or C".

[0147] The term "comprises" or "has" and their synonyms in the present invention are intended to mean including the items listed thereafter and their equivalents as well as other additional items. Depending on the context, the word "if" in this document can be interpreted to mean "when", "while", or "assuming", etc.

[0148] In the context of the present invention (particularly in the claims), the words "a", "an", "the", and "said" and similar demonstrative pronouns are to be construed to cover both the singular and the plural forms of one and more. Further, the description of a numerical range herein is merely a convenient way of indicating that each individual value within the range is included, and each individual value is incorporated into the range as if it were individually recited herein. Finally, the steps of all methods described herein may be performed in any suitable order, unless otherwise indicated herein or clearly contradicted by the context. The use of examples or exemplary language (such as "such as") provided herein is intended to better illustrate the invention and is not intended to limit the scope of the invention, unless otherwise indicated.

[0149] Throughout this document, various headings and subheadings are used to list items. These are included to enhance readability and to simplify the process of finding and referencing materials. These headings and subheadings are not intended to, and do not, affect the interpretation of the claims or in any way limit the scope of the claims. The specific embodiments shown and described herein are illustrative examples of the invention and are not intended to limit the scope of the invention in any way.

[0150] All references cited herein (including publications, patent applications, and patents, etc.) are hereby incorporated by reference as if each reference were individually and specifically indicated to be incorporated by reference and to cover the entire relevant content of the reference.

[0151] Although the invention has been described in connection with certain embodiments and implementations, it is to be understood that the invention is not limited to the embodiments disclosed herein. The disclosure of the invention is intended to cover various variations and equivalent arrangements within the scope of the claims, which scope should be given the broadest interpretation to cover all such variations and equivalent arrangements as are permitted by law.

Claims

1. A bitrate allocation method in video communication, comprising: Encoding a video stream with an encoder, where the video stream is divided into slices associated with the viewpoints of the video stream, and encoding the video stream includes allocating a bitrate for each of the slices; Determining an image quality score for slice rendering in a first iteration, where slice rendering includes decoding and displaying the slice in the first iteration in the viewpoint of the video stream at the allocated bitrate; And After determining the image quality score, adjusting the bitrate allocated to one or more of the slices such that slice rendering in a second iteration after the first iteration includes decoding and displaying the slice at the adjusted bitrate.

2. The method according to claim 1, wherein Performing the bitrate adjustment among the slices for which bitrates have been allocated.

3. The method according to claim 1, wherein Determining the image quality score for slice rendering includes: determining an average image quality score of the slice in the first iteration, and adjusting the bitrate allocated to one or more of the slices, where the adjustment includes modifying the bitrate allocated to one or more of the slices such that the image quality score of each slice after modifying the bitrate converges to the average image quality score.

4. The method according to claim 1, wherein The method further includes: Determining an image quality score for slice rendering in a second iteration; Determining an image quality gain, including determining a change between the image quality score of slice rendering in the first iteration and the image quality score of slice rendering in the second iteration; and Determining whether the image quality gain exceeds an image quality gain threshold.

5. The method according to claim 4, characterized in that Slice rendering in the second iteration is performed after slice rendering in the first iteration, and the method further includes: If it is determined that the image quality gain exceeds the image quality gain threshold, establishing an interval time between slice rendering in the second iteration and slice rendering in a third iteration after the second iteration such that the interval time between slice rendering in the second iteration and slice rendering in the third iteration is longer than the interval time between slice rendering in the first iteration and slice rendering in the second iteration.

6. The method according to claim 5, wherein The method further includes: If it is determined that the image quality gain does not exceed the image quality gain threshold, establishing an interval time between slice rendering in the second iteration and slice rendering in a third iteration after the second iteration such that the interval time between slice rendering in the second iteration and slice rendering in the third iteration is equal to the interval time between slice rendering in the first iteration and slice rendering in the second iteration.

7. The method according to claim 3, wherein The method further includes: Determining a bitrate allocation weight, where adjusting the bitrate allocated to one or more of the slices includes: modifying the bitrate allocated to one or more of the slices according to the bitrate allocation weight, and the greater the difference between the image quality score of the slice and the average image quality score, the greater the value of the bitrate allocation weight.

8. The method according to claim 7, wherein The method further includes: Determining the average image quality score of the slices in the second iteration; Determining the difference between the image quality score of each slice in the second iteration and the average image quality score; and Determine whether the difference in the image quality scores of all the slices exceeds an image quality difference threshold.

9. The method according to claim 8, wherein The method further includes: If it is determined that the difference between the image quality scores of one or more of the slices in the second iteration and the average image quality score exceeds the image quality threshold, adjust the bitrate of the one or more slices in the second iteration.

10. A device for bitrate allocation in video communication, comprising: A non-transitory memory; And A processor configured to execute instructions stored in the non-transitory memory to: Encode a video stream, the video stream being divided into slices associated with the viewpoints of the video stream, wherein encoding the video stream includes allocating a bitrate for each slice; Determine the image quality score of slice rendering at the first iteration, wherein the slice rendering includes decoding and displaying the slice in the first iteration in the viewpoint of the video stream at the allocated bitrate; And After determining the image quality score, adjust the bitrate allocated to one or more of the slices such that the rendering of the slices in the second iteration after the first iteration includes decoding and displaying the slices at the adjusted bitrate.

11. The device according to claim 10, wherein Determining the image quality score of the slice rendering includes: determining the average image quality score of the slice at the first iteration, and adjusting the bitrate allocated to one or more of the slices includes making a modification to the bitrate allocated to the slice such that the image quality score of each slice after modifying the bitrate converges to the average image quality score.

12. The device according to claim 10, characterized in that, The processor is further configured to execute instructions stored in the non-transitory memory to: Determine the image quality score of the slices in the second iteration; Determine an image quality gain, including determining the difference between the image quality score of the slices in the second iteration and the image quality score of the slices in the first iteration; And Determine whether the image quality gain exceeds an image quality gain threshold.

13. The device according to claim 12, characterized in that, The slice rendering in the second iteration is performed after the slice rendering in the first iteration, and the processor is further configured to execute instructions stored in the non-transitory memory for: If it is determined that the image quality gain exceeds the image quality gain threshold, establish an interval time between the slice rendering in the second iteration and the slice rendering in the third iteration after the second iteration such that the duration between the slice rendering in the second iteration and the slice rendering in the third iteration is longer than the interval time between the slice rendering in the first iteration and the slice rendering in the second iteration.

14. The device according to claim 12, characterized in that, The processor is further configured to execute instructions stored in the non-transitory memory to: If it is determined that the image quality gain does not exceed the image quality gain threshold, an interval time between the slice rendering in the second iteration and the slice rendering in the third iteration after the second iteration is established such that the interval time between the slice rendering in the second iteration and the slice rendering in the third iteration after the second iteration is equal to the interval time between the slice rendering in the first iteration and the slice rendering in the second iteration.

15. The device according to claim 10, wherein, The processor is further configured to execute instructions stored in the non-transitory memory to: Determine a bitrate allocation weight, wherein adjusting the bitrate allocated to one or more of the slices includes: modifying the bitrate allocated to one or more of the slices according to the bitrate allocation weight.

16. The device according to claim 10, characterized in that, The processor is further configured to execute instructions stored in the non-transitory memory to: Determine the average image quality score of the slices in the second iteration; Determine the difference between the image quality score of each of the slices in the second iteration and the average image quality score; And Determine whether the difference in the image quality scores of all the slices exceeds an image quality difference threshold.

17. A non-transitory computer-readable storage medium configured to store a computer program for bitrate allocation in video communication, the computer program including instructions executable by a processor for: Encoding a video stream that is divided into slices associated with the viewpoints of the video stream, wherein encoding the video stream includes allocating a bitrate to each of the slices; Determining an image quality score of slice rendering at the first iteration, wherein the slice rendering includes decoding and displaying the slices in the first iteration in the viewpoints of the video stream at the allocated bitrate; And After determining the image quality score, adjusting the bitrate allocated to one or more of the slices such that the slice rendering in the second iteration after the first iteration includes decoding and displaying the slices at the adjusted bitrate.

18. The non-transitory computer-readable storage medium according to claim 17, wherein Performing the bitrate adjustment among the slices to which the bitrate has been allocated, wherein determining the image quality score of the slice rendering includes: determining the average image quality score of the slices at the first iteration, and adjusting the bitrate allocated to one or more of the slices includes modifying the bitrate allocated to the slices such that the image quality score of each of the slices after modifying the bitrate converges to the average image quality score.

19. The non-transitory computer-readable storage medium according to claim 17, wherein The computer program further includes instructions executable by the processor to: Determine the image quality score of the slice rendering in the second iteration; Determine an image quality gain, including determining the change between the image quality score of the slice rendering in the first iteration and the image quality score of the slice rendering in the second iteration; And Determine whether the image quality gain exceeds an image quality gain threshold.

20. The non-transitory computer-readable storage medium according to claim 17, wherein The computer program further includes instructions executable by the processor to: Determine a bitrate allocation weight, wherein adjusting the bitrate allocated to one or more of the slices includes: modifying the bitrate allocated to one or more of the slices according to the bitrate allocation weight.