Video processing method, apparatus, electronic device, and storage medium
By cutting panoramic video frames into multiple screens and generating optimized video streams, the method addresses latency and bandwidth issues in VR devices, improving user experience and server efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2024-03-14
- Publication Date
- 2026-04-27
AI Technical Summary
Existing VR devices face challenges in providing panoramic videos with low latency and reduced bandwidth consumption, leading to user discomfort and increased server processing load.
The method involves cutting panoramic video frames into multiple video screens with larger aspect ratios than the playback device, generating two video streams with different GOP lengths, and selecting screens based on playback parameters to reduce latency and bandwidth.
This approach reduces screen display latency to under 25 ms and decreases bandwidth usage, enhancing user experience and server efficiency by allowing local rendering and pre-processing video streams.
Smart Images

Figure 2026513437000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority to the Chinese invention patent application filed on April 3, 2023, titled "Video Processing Method, Apparatus, Electronic Device and Storage Medium," with application number 202310348154.9, the entirety of which is incorporated into this disclosure by reference.
[0002] Technical field This disclosure relates to the field of image processing, and more particularly to video processing methods, apparatus, electronic devices, and storage media. [Background technology]
[0003] With the increasing popularity of VR (Virtual Reality) devices, more and more users are requesting panoramic videos of concerts and other events on their VR devices to experience a more realistic visual experience.
[0004] When providing on-demand panoramic video on VR devices, on the one hand, the head action delay must be controlled to within 25 ms from the time the user's head action occurs until the updated screen is displayed on the VR device. The delay during this period must be less than 25 ms; otherwise, the user will experience noticeable throbbing and even dizziness. On the other hand, the bandwidth used by the server to transmit the panoramic video to the VR device must be reduced, enabling more low-bandwidth users to access on-demand panoramic video on their VR devices. [Overview of the project]
[0005] Embodiments of this disclosure provide video processing methods, apparatus, electronic devices, and storage media that reduce the delay between a user's head action and the display of the updated screen on a VR device when panoramic video is available on demand on a VR device, reduce the bandwidth occupied by the server to transmit the panoramic video to the VR device, and improve the user's experience of watching panoramic video on demand on a VR device.
[0006] In the first aspect, embodiments of the present disclosure are as follows: The process involves obtaining a video stream awaiting processing, which includes multiple video frames, wherein the aspect ratio of each video frame among the multiple video frames is greater than the aspect ratio of the video playback device, and cutting each video frame among the multiple video frames into multiple video screens, wherein the aspect ratio of each video screen is greater than the aspect ratio of the video playback device. For video screens among the plurality of video screens where the position of the center point of the screen is the same, a first video screen stream and a second video screen stream are generated in the time order of the playback timestamps of the video screens, wherein the first screen group GOP of the first video screen stream is a standard GOP, and the length of the second screen group GOP of the second video screen stream is smaller than the length of the first screen group GOP of the first video screen stream. The present invention provides a video processing method that includes determining a video screen to be played from the first video screen stream and the second video screen stream based on the video playback parameters of the video playback device.
[0007] In a second aspect, embodiments of the present disclosure are as follows: A video acquisition unit for acquiring a video stream awaiting processing, which includes multiple video frames, wherein the aspect ratio of each video frame among the multiple video frames is greater than the aspect ratio of a video playback device, A video cutting unit for cutting each of the plurality of video frames into a plurality of video screens, wherein the field of view of the video screen is larger than the field of view of the video playback device, A video generation unit for generating a first video screen stream and a second video screen stream in the time order of the playback timestamps of multiple video screens that have the same screen center point position, wherein the first screen group GOP of the first video screen stream is a standard GOP, and the length of the second screen group GOP of the second video screen stream is smaller than the length of the first screen group GOP of the first video screen stream, The present invention provides a video processing device that includes a video distribution unit for determining a video screen awaiting playback from the first video screen stream and the second video screen stream based on the video playback parameters of the video playback device.
[0008] In a third aspect, an embodiment of the present disclosure provides an electronic device including a processor and a memory arranged to store computer executable instructions, wherein when the computer executable instructions are executed, the memory causes the processor to perform the steps of the method described in the first aspect.
[0009] In a fourth aspect, embodiments of the present disclosure provide a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, realize steps of the method described in the first aspect.
[0010] In one or more embodiments of the present disclosure, a video stream awaiting processing is acquired, which includes a plurality of video frames having a field of view larger than the field of view of a video playback device; each video frame of the plurality of video frames is cut into a plurality of video screens; for video screens whose field of view is larger than the field of view of the video playback device and whose screen center point position is the same among the plurality of video screens, a first video screen stream and a second video screen stream are generated in the time order of the playback timestamps of the video screens; the first screen group GOP of the first video screen stream is a standard GOP; the length of the second screen group GOP of the second video screen stream is smaller than the length of the first screen group GOP of the first video screen stream; and a video screen awaiting playback is determined from the first video screen stream and the second video screen stream based on the video playback parameters of the video playback device.
[0011] According to this embodiment, on the one hand, an appropriate waiting video screen to be played on the user's video playback device is selected from the first video screen stream and the second video screen stream. Since the field of view of the waiting video screen is larger than the field of view of the user's video playback device, when the user views the waiting video screen via the VR device, it is not necessary to transmit a panoramic 360-degree video screen to the user, and it is possible to satisfy the user's small range of head movement. When panoramic video on demand on a VR device, the bandwidth occupied by the server to transmit the panoramic video to the VR device is reduced, improving the user's experience of watching panoramic video on demand on a VR device. On the other hand, the operation to generate the first and second video screen streams can be performed before video playback. During video playback, the appropriate waiting video screen from the first and second video screen streams is selected based on the video playback parameters of the user's video playback device, and the video playback device is instructed to do so. This significantly reduces the processing load on the server and improves the server's processing efficiency. As a result, when performing panoramic video on demand in a VR device, the delay from when the user's head action occurs until the updated screen is displayed on the VR device is reduced, improving the user's experience of watching panoramic video on demand in a VR device.
[0012] To better illustrate one or more embodiments of this disclosure or the technical concepts in the prior art, the drawings that may be used in the description of the embodiments or prior art are briefly described below. Obviously, the drawings in the following description are only a few embodiments of the disclosure, and those skilled in the art can obtain other drawings based on these without any creative effort. [Brief explanation of the drawing]
[0013] [Figure 1] This is a schematic flowchart of a video processing method provided in one embodiment of the present disclosure. [Figure 2]It is a schematic configuration diagram of a video processing system provided in an embodiment of the present disclosure. [Figure 3] It is a schematic diagram of a cut video screen provided in an embodiment of the present disclosure. [Figure 4] It is a schematic diagram of an angle-of-view interval provided in an embodiment of the present disclosure. [Figure 5] It is a schematic diagram of a video screen obtained by cutting a video frame in a video stream provided in an embodiment of the present disclosure. [Figure 6] It is a schematic diagram of an IPP flow and an IP flow provided in an embodiment of the present disclosure. [Figure 7] It is a schematic flowchart of selecting a video screen waiting for playback provided in an embodiment of the present disclosure. [Figure 8] It is a schematic diagram of a user-switching video screen provided in an embodiment of the present disclosure. [Figure 9] It is a schematic scene diagram of selecting a video screen waiting for playback provided in an embodiment of the present disclosure. [Figure 10] It is a schematic configuration diagram of a video processing apparatus provided in an embodiment of the present disclosure. [Figure 11] It is a schematic configuration diagram of an electronic device provided in an embodiment of the present disclosure.
Embodiments for Carrying Out the Invention
[0014] In order for those skilled in the art to better understand the technical solutions in one or more embodiments of the present disclosure, hereinafter, referring to the drawings in one or more embodiments of the present disclosure, the technical solutions in one or more embodiments of the present disclosure will be clearly and completely described. Obviously, the described embodiments are only some embodiments of the present disclosure, not all embodiments. Based on one or more embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative labor shall fall within the protection scope of the present disclosure.
[0015] Embodiments of this disclosure provide a video processing method that reduces the delay between a user's head action and the display of the updated screen on a VR device when panoramic video is available on demand on a VR device, reduces the bandwidth occupied by the server to transmit the panoramic video to the VR device, and improves the user's experience of viewing panoramic video on demand on a VR device. The video processing method provided by one embodiment of this disclosure can be executed by a background server that can be located in an edge cloud machine room, and by implementing the video processing method in an edge cloud machine room, the delay when a user views on-demand video via a VR device can be reduced.
[0016] Figure 1 is a schematic flowchart of a video processing method provided by one embodiment of the present disclosure. As shown in Figure 1, the process is: Step S102 involves acquiring a video stream awaiting processing, which includes multiple video frames having a field of view larger than the field of view of the video playback device. Step S104 involves cutting each of the multiple video frames into multiple video screens so that the aspect ratio of the video screen is greater than the aspect ratio of the video playback device. Step S106: For video screens among multiple video screens that have the same screen center point position, a first video screen stream and a second video screen stream are generated in the time order of the playback timestamps of the video screens, the first screen group GOP of the first video screen stream is a standard GOP, and the length of the second screen group GOP of the second video screen stream is less than the length of the first screen group GOP of the first video screen stream. The process includes step S108, which determines a video screen to be played from a first video screen stream and a second video screen stream based on the video playback parameters of a video playback device.
[0017] According to this embodiment, on the one hand, an appropriate waiting video screen to be played on the user's video playback device is selected from the first video screen stream and the second video screen stream. Since the field of view of the waiting video screen is larger than the field of view of the user's video playback device, when the user views the waiting video screen via the VR device, it is not necessary to transmit a panoramic 360-degree video screen to the user, and it is possible to satisfy the user's small range of head movement. When panoramic video on demand on a VR device, the bandwidth occupied by the server to transmit the panoramic video to the VR device is reduced, improving the user's experience of watching panoramic video on demand on a VR device. On the other hand, the operation to generate the first and second video screen streams can be performed before video playback. During video playback, the appropriate waiting video screen from the first and second video screen streams is selected based on the video playback parameters of the user's video playback device, and the video playback device is instructed to do so. This significantly reduces the processing load on the server and improves the server's processing efficiency. As a result, when performing panoramic video on demand in a VR device, the delay from when the user's head action occurs until the updated screen is displayed on the VR device is reduced, improving the user's experience of watching panoramic video on demand in a VR device.
[0018] Step S104 in Figure 1 can be abbreviated as the video screen cutting step, the process of obtaining the first video screen stream and the second video screen stream in step S106 can be collectively referred to as the video stream conversion step, and in this embodiment, the process of storing the first video screen stream and the second video screen stream can be further included and can be abbreviated as the video stream storage step. In a video-on-demand scene, the process of selecting a video screen waiting to be played that matches the video playback parameters of the video playback device from the stored first video screen stream and the second video screen stream in step S108, and sending it to the video playback device for playback, can be abbreviated as the video distribution step.
[0019] Based on this, one embodiment of the present disclosure further provides a video processing system. Figure 2 is a schematic diagram of the configuration of the video processing system provided by one embodiment of the present disclosure. As shown in Figure 2, the system is located in an edge cloud room and includes an offline transcoding service, a storage service, and a real-time distribution service, the offline transcoding service including a video screen cutting service and a video stream conversion service. In Figure 2, the video screen cutting service is used to implement the video screen cutting step, the video stream coding service is used to implement the video stream conversion step, the storage service is used to implement the video stream storage step, and the real-time distribution service is used to implement the video distribution step. As can be seen from Figure 2, after acquiring an on-demand video stream, the edge cloud room stores a first video screen stream and a second video screen stream through the process in Figure 1, acquires the video playback parameters of the user's video playback device, and then delivers a video screen waiting to be played that matches the video playback parameters to the user's video playback device for playback. The user's video playback device may be a VR device. The on-demand video stream (also called a video source) may be a panoramic 360-degree video stream.
[0020] According to the system in Figure 2, before a user watches on-demand video via a VR device, the system can pre-process and obtain and store a first video screen stream and a second video screen stream. When the user watches on-demand video via the VR device, it is only necessary to deliver the video screen waiting to play to the user via a real-time distribution service. This allows for high concurrency, and different users can watch different on-demand videos from different viewpoints via the real-time distribution service.
[0021] High concurrency, as used here, means that multiple servers can be set up to realize a real-time streaming service, and while scheduling allows multiple servers to support high concurrency, the real-time streaming service only needs to deliver a video screen waiting to play to the user, requiring less computing power. Therefore, the concurrency of individual servers can be increased by reducing the computing power of each server and the downstream bandwidth occupied by the video screen waiting to play. Furthermore, increasing the concurrency of a single server by reducing the computing power of a single server and the downstream bandwidth occupied by the video screen waiting to play also has the effect of reducing server usage costs.
[0022] The process shown in Figure 1 will be explained in detail below.
[0023] In step S102 above, a video stream awaiting processing is acquired. The video stream awaiting processing is an on-demand video stream, which includes multiple video frames, and in one embodiment, each of the multiple video frames is a panoramic video frame, i.e., a panoramic 360-degree video frame, and accordingly, the video stream is a panoramic 360-degree video stream. In other embodiments, the video stream may be a wide-angle video stream of other angles, such as a 270-degree or 180-degree video stream, where the field of view of each of the multiple video frames in the video stream awaiting processing can be preset to be greater than the field of view of the video playback device. For example, if the field of view of the video playback device is 90 degrees, then each of the multiple video frames in the video stream awaiting processing is greater than 90 degrees, and a video stream in which such a field of view is greater than the field of view of the video playback device is designated as the video stream awaiting processing.
[0024] In step S104 described above, each of the multiple video frames is cut into multiple video screens, and the field of view of each video screen obtained by the cut is larger than the field of view of the user's video playback device.
[0025] In one embodiment, the video screen obtained from a cut is called VAM (Visible Area + Margin), and if the video playback device is a VR device, the field of view of the video playback device can be represented by the FOV (Field of View) of the VR device. The field of view corresponding to VAM is larger than the field of view of the video playback device, and in the VAM screen, the field of view of the visible area is equal to the field of view of the video playback device, and the margin / allowance is the field of view in which more of the VAM screen appears relative to the video playback device. Figure 3 is a schematic diagram of a video screen obtained from a cut provided by one embodiment of this disclosure, and as shown in Figure 3, the field of view of this video screen VAM is larger than the field of view of the VR device.
[0026] Setting the field of view of the video screen obtained by cutting to be larger than the field of view of the VR device video playback device, and transmitting the video screen to the VR device for video on demand, has two advantages: 1. Since the video screen is cut from the video frame, the resolution of the video screen is smaller than the resolution of the video frame. By using the video screen instead of the video frame to transmit to the VR device for on-demand use, the code rate of network transmission is effectively reduced. 2. When a user wears a VR device and watches on-demand video with a small range of head movement, the VR device can directly use the margin / margin area in the video screen to perform local rendering, satisfying the user's small range of head movement. Since there is no need to acquire a new video screen from the backend, the screen update delay can be controlled to a preset time, for example, within 25 ms, and the user can avoid feeling screen sluggish or dizziness.
[0027] In one embodiment, cutting each video frame from a plurality of video frames into multiple video screens is, Obtain video screen cut parameters including the resolution of the video screen, the scaling ratio of the video screen to the video frame, the margin of the field of view of the video screen relative to the field of view of the video playback device, and the field of view and viewing angle of the video playback device between adjacent video screens. This includes cutting each video frame out of multiple video frames into multiple video screens based on video screen cut parameters.
[0028] Specifically, the video screen cut parameters include the resolution of the video screen, the scaling ratio of the video screen to the video frame, the field of view of the video playback device, the margin of the video screen's field of view relative to the field of view of the video playback device, and the field of view interval between two adjacent video screens. These parameters allow for precise cutting of video frames within a video stream, resulting in a usable video screen.
[0029] Thus, in this embodiment, several parameters such as the resolution of the video screen, the scaling ratio of the video screen to the video frame, the margin of the video screen's field of view relative to the field of view of the video playback device, and the field of view and viewing angle of the video playback device between adjacent video screens can be determined as video screen cut parameters. By doing so, video frames can be precisely cut based on the video screen cut parameters, and different video screens can be obtained.
[0030] In one embodiment, the resolution of the video screen, the scaling ratio of the video screen to the video frame, and the margin of the video screen's field of view relative to the field of view of the video playback device are as follows: The steps include obtaining a pre-set range of values for the resolution of the video screen (first value), a range of values for the scaling ratio (second value), and a range of values for the field of view margin (third value), A constraint used to express that the ratio result of the video screen resolution and scaling ratio is equal to the above-mentioned field of view margin, by subtracting the field of view of the video playback device from the result of the ratio of the video screen resolution and scaling ratio, comprising the steps of obtaining constraints on the video screen resolution, scaling ratio, and field of view margin, The video screen resolution, scaling ratio, and field of view margin that satisfy the constraints are determined by a step that determines them within the ranges of a first, second, and third value, respectively.
[0031] First, obtain the first value range of the pre-set video screen resolution. The resolution is Crop pitch and Crop yaw These can be expressed as , representing the resolution in the pitch and yaw directions of the video screen, respectively. Higher resolution results in a sharper video screen, but a larger video screen size also increases the transmission code rate. The range of the first value can represent the maximum and minimum resolution in the pitch direction, and the maximum and minimum resolution in the yaw direction.
[0032] Next, we obtain the range of a pre-set second value for the video screen relative to the video frame scaling ratio. The scaling ratio is denoted as Scale_factor, and it scales equally in both the yaw and pitch directions. The range of the second value can be (0,1), and a larger value indicates a lower degree of scaling, meaning that at the same resolution, the video screen covers less screen information. Conversely, a smaller value indicates a higher degree of scaling, meaning that at the same resolution, the video screen covers more screen information, but the image clarity is lower. A scaling ratio of 1 indicates that the video screen is not scaled. Multiple scaling ratio values can be set to match different VR device applications, satisfying the user's need to view high-resolution screens while also satisfying their need to check more screen information.
[0033] Next, the range of a third value for the field of view of a video screen relative to the field of view margin of the video playback device, i.e., the range of the third value of the margin / surplus area Margin, is obtained. Margin is distinguished in the yaw and pitch directions, and the larger the Margin value, the greater the network delay that can be covered, the less likely VAM (video image transitions) to occur, and the fewer VAMs (video images) there are, resulting in less memory pressure, but the transmission code rate is increased and the clarity at the same resolution decreases. pitch and Margin yaw It is divided into two sections, each displaying the size of the frame in the pitch direction and the size of the frame in the yaw direction.
[0034] Then, constraints on resolution, scaling ratio, and field of view margin are obtained. The constraints are used to show that the ratio result of resolution and scaling ratio is equal to the field of view margin by subtracting the field of view of the video playback device. In this case, the field of view (FOV) of the video playback device is a fixed value, and the FOV was determined after the model number of the video playback device was determined. FOV is FOV pitch and FOV yaw These include the pitch angle and yaw angle, respectively. The constraints are expressed by the following equations (1), (2), (3), and (4).
[0035]
number
[0036] In one embodiment, the scaling ratio can first be set to two values, 1 and 0.3. Of course, the scaling ratio can be a fixed value, can be adjusted by the user on the VR device, or can be automatically switched to different values depending on the network conditions in the background. Next, the resolution value and the field of view margin value are determined based on the range of the first value, the range of the third value, and constraint equations (1) to (4) described above. The optimal value for the field of view margin is approximately 35 degrees.
[0037] Thus, according to this embodiment, the resolution of the video screen, the scaling ratio of the video screen to the video frame, and the field of view margin of the video screen to the field of view of the video playback device can be accurately determined based on the range of a first value for the resolution, the range of a second value for the scaling ratio, the range of a third value for the field of view margin, and constraints on the resolution, scaling ratio, and field of view margin.
[0038] In one embodiment, the viewing angle interval is determined by a step that determines the viewing angle interval between adjacent video screens based on the viewing angle margin and the viewing angle of each video frame. In this embodiment, the viewing angle interval between adjacent video screens can be accurately determined based on the viewing angle margin and the viewing angle of each video frame, and furthermore, the video screen cut parameters can be accurately determined.
[0039] In one embodiment, after obtaining the resolution of the video screen, the scaling ratio of the video screen to the video frame, the field of view margin of the video screen relative to the field of view of the video playback device, and the field of view of the video playback device, the field of view interval between adjacent video screens is determined based on the field of view margin and the field of view of each video frame. Specifically: We obtain a first constraint on the viewing angle interval, which means that the viewing angle interval is less than or equal to the field of view margin. We obtain a second constraint on the viewing angle interval, which shows that the field of view of each video frame is divisible by the viewing angle interval. Determine the viewing angle interval based on the first constraint condition of the viewing angle interval and the second constraint condition of the viewing angle interval.
[0040] FIG. 4 is a schematic diagram of the viewing angle interval provided by an embodiment of the present disclosure. As shown in FIG. 4, the viewing angle interval between two adjacent video screens VAM can be denoted as Rotation. FIG. 4 also shows the margin / surplus area Margin / 2 of the right video screen VAM.
[0041] First, obtain the first constraint condition for the viewing angle interval. The first constraint condition indicates that the viewing angle interval is less than or equal to the angle-of-view margin. The first constraint condition for the viewing angle interval Rotation can be expressed by the following formula (5).
[0042]
Equation
[0043] Next, obtain the second constraint condition for the viewing angle interval, which indicates that the angle of view of each video frame is divisible by the viewing angle interval. The second constraint condition for the viewing angle interval Rotation can be expressed by the following formulas (6) and (7).
[0044]
Equation
[0045] Finally, the visual angle interval is determined based on the first constraint and the second constraint on the visual angle interval.
[0046] In this embodiment, by setting the viewing angle interval to be less than or equal to the field of view margin, it is possible to ensure that each video screen cut according to the viewing angle interval always covers the field of view (FOV) of the video playback device. Specifically, if the user's head action range in the VR device is within the field of view margin, the screen the user needs can be obtained by rendering the currently streamed video screen. However, if it exceeds the field of view margin, it is necessary to stream a video screen with a viewing angle adjacent to the user. As shown in Figure 4, since the viewing angle interval between the video screen with an adjacent viewing angle and the currently streamed video screen is less than or equal to the field of view margin, the screen the user wants to see after a head action can be obtained by rendering from the adjacent video screen. As a result, no matter how the user moves their head, the necessary screen can be obtained by rendering from the currently streamed video screen or a video screen with an adjacent viewing angle, eliminating the need to stitch together different video screens and improving the efficiency of screen display in the VR device.
[0047] In this embodiment, by arranging the video frames so that their aspect ratios are divisible by the viewing angle interval, it is possible to ensure the smallest possible overlap between video screens (VAMs) and to ensure that the center points of the video screens (VAMs) are distributed as uniformly as possible. Each video screen cut along the viewing angle interval covers the entire aspect ratio of the video frame.
[0048] Therefore, according to this embodiment, based on the first and second constraints described above, the resulting viewing angle interval can be ensured so that no matter how the user moves their head, the necessary screen can be obtained by rendering from the currently streamed video screen or the video screen of the adjacent viewing angle, eliminating the need to stitch together different video screens, improving the efficiency of screen display for VR devices, ensuring the smallest possible overlap between video screens (VAMs), and ensuring that the screen center points of the video screens (VAMs) are distributed as uniformly as possible.
[0049] In one embodiment, the position value of the center point of all video screens is calculated based on the viewing angle interval. For example, if the viewing angle interval is 20 degrees in both the pitch and yaw directions, the center point of each video screen is calculated in the order of (0,0), (0,0), (0,20), (0,40)...(20,20), (20,40)... and so on. In the case of a panoramic video stream, the video screen VAM is Rotation pitch and Rotation yaw Each VAM is cut within the ranges of [0,360] and [0,180]. Due to the characteristics of ERP (Equirectangular Projection) images, one VAM should be cut at pitch=0 and pitch=180, which are the north and south poles. Therefore, the total number of VAMs in the video screen is
[0050]
number
[0051] That is the case.
[0052] In this embodiment, the resolution of the video screen determined above, the scaling ratio of the video screen to the video frame, the field of view margin of the video screen relative to the field of view of the video playback device, and the field of view of the video playback device and the viewing angle between adjacent video screens are determined as video screen cut parameters.
[0053] In one embodiment, cutting each video frame from a plurality of video frames into multiple video screens based on video screen cut parameters specifically means: The aspect ratio of the video screen is determined based on the aspect ratio and aspect ratio margin of the video playback device. For each of the multiple video frames in the video stream, the video frame is mapped to a unit sphere, and the mapped image is scaled based on the scaling ratio. For each of the multiple video frames in a video stream, the scaled image of that video frame is cropped based on the field of view, resolution, and viewing angle interval of the video screen, thereby obtaining the video screen at each viewing angle of that video frame.
[0054] First, the field of view of the video screen (VAM), i.e., the size of the VAM in Figure 3, is determined based on the field of view (FOV) and the field of view margin (Margin). As can be seen from Figure 3, the size of the VAM is equal to the FOV plus the field of view margin (Margin). Next, for each video frame in the video stream, the video frame is mapped to a unit sphere, and the mapped image is scaled based on the scaling ratio. If there are multiple scaling ratio values, scaling must be performed once for each scaling ratio value. Finally, for each video frame, the scaled image of the video frame is cropped based on the field of view size of the video screen (VAM), the resolution, and the viewing angle interval, to obtain the video screen at each viewing angle of the video frame.
[0055] In this embodiment, the video frame is first mapped to a unit sphere, the mapped image is scaled based on the scaling ratio, and then the scaled image is cut based on the field of view, resolution, and viewing angle interval of the video screen to obtain the video screen at each viewing angle. By cutting the scaled image, it is possible to ensure that the video screen seen by the user is the scaled image, thereby satisfying the user's need for scaling.
[0056] In one embodiment, based on the field of view, resolution, and viewing angle interval of the video screen, the scaled image of the video frame is cropped to obtain the video screen at each viewing angle of the video frame, which specifically involves: The center point of the video screen is aligned with the center point of the scaled image of the video frame, and the scaled image of the video frame is cropped according to the resolution and the field of view of the video screen to obtain the video screen at the reference viewing angle of the video frame. Based on the viewing angle interval, the center point of the video screen is moved onto the scaled image of the video frame, and the scaled image of the video frame is cropped according to the resolution and the field of view of the moved video screen to obtain the video screen at other viewing angles of the video frame.
[0057] First, the center point of the video screen is aligned with the center point of the scaled image of the video frame. Then, according to a predetermined resolution and the field of view of the video screen, the scaled image of the video frame is cropped to obtain an image whose field of view is equal to the field of view of the video screen and whose resolution is equal to the predetermined resolution. This image is the video screen at the reference viewing angle of the video frame.
[0058] Then, the center point of the video screen is moved onto the scaled image of this video frame according to the viewing angle interval, and the previous cutting operation is repeated. Cuts are made on the scaled image of the video frame according to the resolution and the field of view of the moved video screen, and an image is obtained in which the field of view is equal to the field of view of the video screen and the resolution is equal to a predetermined resolution. This image is the video screen at a different viewing angle of the video frame. Each time the center point of the video screen is moved once according to the viewing angle interval, a cut is made and a video screen at a different viewing angle is obtained. The center point of the video screen is moved several times according to the viewing angle interval until all the image content of the scaled image is cut.
[0059] Through the above process, any video frame can be cut to obtain the video screen for each viewing angle. In another embodiment, the scaled image can be moved according to the viewing angle interval, and the video screen for other viewing angles can be obtained by cutting within the moved scaled image.
[0060] In one embodiment, the center point of the video screen is aligned with the center point of the scaled image, the position of the center point of the video screen is determined to be (0,0), the center point of the video screen is moved on the scaled image according to the viewing angle interval, the position of the moved center point is determined based on the viewing angle interval and the initial position (0,0), and the position of the center point of each video screen obtained in the cut is determined based on the viewing angle interval.
[0061] In another embodiment, after the viewing angle interval is determined, the screen center point of each video screen is also determined. Specifically, the coordinates of the screen center points of adjacent video screens can be determined sequentially from (0,0) according to the viewing angle intervals in different directions. For example, if the viewing angle interval in different directions is 20 degrees, the screen center points of each video screen will be (0,0), (0,20), (0,40)...(20,20), (20,40)... and so on.
[0062] As can be seen from this, the position of the center point of the video screen, determined based on the viewing angle interval, can represent the viewing angle information of the video screen. Through the above process, multiple video screens at different viewing angles can be obtained by cutting any video frame, and the position of the center point of the corresponding video screen will differ for video screens at different viewing angles.
[0063] In this embodiment, the field of view of the video screen is proportional to the network delay, the degree of scaling of the video frame mapped to a unit sphere determines the clarity of the video screen, and the resolution of the video screen determines the size of the video screen. According to this embodiment, it is possible to obtain video screens at different viewing angles by precisely cutting the video frame based on the viewing angle interval, resulting in high cutting efficiency.
[0064] Figure 5 is a schematic diagram of video screens obtained by cutting video frames in a video stream provided by one embodiment of the present disclosure. As shown in Figure 5, for each video frame in the video stream, multiple video screens are obtained by cutting, and each video screen corresponds one-to-one with each viewing angle. n and m are positive integers of 1 or more.
[0065] In the process shown in Figure 1, after obtaining each video screen for each video frame, step S106 is executed to generate a first video screen stream and a second video screen stream in the time order of the playback timestamps of the video screens for which the center point of multiple video screens is the same. That is, for video screens with the same viewing angle among multiple video screens, such as the video screens with the same viewing angle in Figure 5, a first video screen stream and a second video screen stream are generated in the time order of the playback timestamps of the video screens. The first screen group GOP of the first video screen stream is a standard GOP, and the length of the second screen group GOP of the second video screen stream is smaller than the length of the first screen group GOP of the first video screen stream.
[0066] In one embodiment, for video screens with the same center point position among multiple video screens, a first video screen stream and a second video screen stream are generated in the time order of the playback timestamps of those video screens. For multiple video screens where the center point of the screen is located, the initial video screen stream is obtained by arranging the video screens in the time order of their playback timestamps. The initial video screen stream is converted into a first video screen stream and a second video stream, wherein the first screen group GOP of the first video screen stream contains one I-frame and at least two P-frames, and the second screen group GOP of the second video screen stream contains one I-frame and at most two P-frames.
[0067] Specifically, first, for each video screen in each video frame, video screens corresponding to the same viewing angle, i.e., video screens with the same center point position, are arranged in chronological order of the playback timestamps to obtain initial video screen streams for each viewing angle. Then, these initial video screen streams for each viewing angle are encoded into a first video screen stream and a second video screen stream, respectively.
[0068] For example, if there are 10 video frames, and each video frame is cut to obtain 30 video screens for different viewing angles, and the 30 corresponding viewing angles for each video frame are the same, then the 10 video screens for the first viewing angle (from each video frame) are arranged in chronological order of the playback timestamps to obtain the initial video screen stream for the first viewing angle, and the 10 video screens for the second viewing angle are arranged to obtain the initial video screen stream for the second viewing angle. Similarly, after obtaining the initial video screen streams for all 30 viewing angles, each initial video screen stream for each viewing angle is encoded into the first video screen stream and the second video screen stream, respectively.
[0069] A GOP (Group of Pictures, also called a picture group or screen group) is a set of consecutive images within an MPEG-encoded movie or view stream, beginning with an I-frame and ending with the next I-frame. I-frames (Intra-Coded Pictures): Also called intra-coded frames, these are keyframes, independent frames containing all the information, and can be decoded independently without referencing other images, easily understood as static screens. P-frames (Predictive Coded Pictures): Also called inter-predictive coded frames, these must be coded by referencing the previous I-frame. They show the difference between the current frame and the previous frame (which may or may not be an I-frame or a P-frame). Decoding requires overlaying the previously cached screen with the difference defined in that frame to generate the final screen. While P-frames typically occupy fewer data bits than I-frames, their complex dependencies on previous P and I reference frames make them highly sensitive to transmission errors. B-frames (Bidirectionally Predictive Coded Pictures): Also called bidirectional predictive coded frames, B-frames record the difference between the current frame and the preceding and succeeding frames. In other words, to decrypt a B-frame, it is necessary not only to retrieve the previous cached screen, but also to decrypt the subsequent screen and obtain the final screen by overlapping the main frame data of the preceding and succeeding screens. B-frames have a high compression ratio, but the decryption performance requirements are high.
[0070] In this embodiment, the first screen group GOP of the first video screen stream is a standard GOP, which is any GOP compatible with existing encoding standards (e.g., H264, H265, etc.), for example, a GOP in which one I-frame is followed by multiple P-frames. The length of the second screen group GOP of the second video screen stream is less than the length of the first screen group GOP of the first video screen stream, i.e., the number of frames contained in one GOP in the second video screen stream is less than the number of frames contained in one GOP in the first video screen stream. In one case, the first screen group GOP of the first video screen stream contains one I-frame and at least two P-frames, and the second screen group GOP of the second video screen stream contains one I-frame and at most two P-frames.
[0071] The frame structure of the first video screen stream can be called an IPP frame structure, representing that one GPO contains one I-frame and at least two P-frames. The frame structure of the second video screen stream is called an IP frame structure, representing that one GOP contains one I-frame and at most two P-frames. After obtaining the initial video screen stream for each viewing angle, each video screen in the initial video screen stream for each viewing angle is encoded to obtain a video screen with an IPP frame structure, and each video screen in each video screen stream is encoded to obtain a video screen with an IP frame structure. That is, each video screen in the initial video screen stream for each viewing angle is encoded into an I-frame or a P-frame. After encoding, the number of video screens in the first video screen stream is equal to the number of video screens in the second video screen stream, and equal to the number of video screens in the video screen stream for each viewing angle.
[0072] In this embodiment, in the case of an IPP frame structure, one GOP includes one I-frame and at least two P-frames, and in the case of an IP frame structure, one GOP includes one I-frame and at most two P-frames. Furthermore, neither the IPP frame structure nor the IP frame structure includes B-frames, and in both the IPP frame structure and the IP frame structure, each video screen is encoded by referencing only the previous video screen.
[0073] In this embodiment, by converting the initial video screen stream for each viewing angle into an IPP stream (first video screen stream) and an IP stream (second video screen stream), the video screen can be delivered to the user at the frame level when the user watches on-demand video, enabling switching of viewing angles. This reduces the computational pressure on the server to deliver the video screen, reduces screen latency, and improves the user experience.
[0074] Figure 6 is a schematic diagram of the IPP flow and IP flow provided in one embodiment of the present disclosure. As shown in Figure 6, each viewing angle has one initial video screen stream, and the initial video screen stream of each viewing angle can be encoded into an IPP flow and an IP flow, respectively, thereby obtaining one IPP flow and one IP flow under each viewing angle. K and T are positive integers greater than or equal to 1.
[0075] In one embodiment, the above method generates first and second video screen streams, For a first video screen stream and a second video screen stream generated by video screens with the same center point position among multiple video screens, a name is assigned to each video screen in the first video screen stream and each video screen in the second video screen stream according to the center point position, the identifier of the first screen group GOP, the identifier of the second screen group GOP, and the video screen number setting rule, and name information is obtained. The method further includes storing the first and second video screen streams based on the naming information of each video screen in the first video screen stream and the naming information of each video screen in the second video screen stream.
[0076] Specifically, for each video screen in the IPP flow for each viewing angle, a first video screen stream is generated by video screens with the same screen center point position among multiple video screens. The name is assigned to each video screen in the IPP flow for each viewing angle according to the screen center point position of the video screen, the identifier of the first screen group GOP, such as the IPP frame structure identifier, and a pre-configured video screen number setting rule. The resulting name information includes the screen center point position of the video screen, the IPP frame structure identifier, and the video screen number. The video screen number setting rule allows the video screen number to be set from 0.
[0077] Similarly, for a second video screen stream generated by multiple video screens with the same screen center point position, i.e., for each video screen in the IP flow of each viewing angle, a name is assigned to each video screen in the IP flow of each viewing angle according to the screen center point position of the video screen, the identifier of the second screen group GOP, such as the IP frame structure identifier, and a pre-configured video screen number setting rule. The resulting name information includes the screen center point position of the video screen, the IP frame structure identifier, and the video screen number, and the video screen number setting rule allows the video screen number to be set from 0.
[0078] Next, the first video screen stream and the second video screen stream are stored based on the name information of each video screen in the first video screen stream and the name information of each video screen in the second video screen stream.
[0079] Thus, according to this embodiment, the names of each video screen in the first video screen stream and the second video screen stream can be determined according to the screen center point position, the identifier of the first screen group GOP, the identifier of the second screen group GOP, and the video screen number setting rules. As a result, the screen center point position, frame structure, and number of each video screen can be determined from the name, and each video screen can be easily distinguished.
[0080] In one specific embodiment, the process of steps S102 to S106 is completed before the user watches on-demand video, and the first video screen stream and the second video screen stream are stored in NFS. When the process of steps S102 to S106 is executed by the first server, in a video on-demand scene, the first server may select a video screen waiting to be played from the stored first video screen stream and the second video screen stream and send it to the video playback device, or another server other than the first server may select a video screen waiting to be played from the stored first video screen stream and the second video screen stream and send it to the video playback device.
[0081] If there are multiple video screen values for the scaling ratio of a video frame, the initial video screen stream at the same screen center point position contains multiple substreams, each corresponding to a scaling ratio value, and each substream is converted into a first video screen stream and a second video screen stream, respectively. In this case, when storing the first and second video screen streams in NFS, the following directory structure can be used to quickly find the required video screen. The target structure includes: Layer 1: setting up a total folder representing the entry point of the video screen streams; Layer 2: naming the folders using the screen center point position, scaling ratio value, and frame structure identifier, where the screen center point position includes pitch and yaw angle coordinates; and Layer 3: storing the corresponding video screen in each folder according to the video screen number.
[0082] Thus, in the example above, if there are multiple scaling ratio values, the first and second video screen streams can be stored according to the scaling ratio values.
[0083] After obtaining the first and second video screen streams, step S108 obtains the video playback parameters of the video playback device, determines the video screen to be played from the first and second video screen streams based on the video playback parameters of the video playback device, and distributes it to the video playback device.
[0084] In one embodiment, when each panoramic video frame in a panoramic video stream is cut into multiple video screens, each video screen corresponds to a single screen center point position, i.e., a viewing angle, for each panoramic video frame, and the viewing angle of each video screen can cover all the contents of the panoramic video frame. When generating and storing the first and second video screen streams for each viewing angle, the first video screen stream is stored in H265 format and the second video screen stream is stored in H264 format. When distributing the video screens, there is no need to transcode them again, and they can be distributed directly, saving server time and computing power, and supporting simultaneous access to the distribution service. The server performing the distribution can adopt an ARM architecture, thereby further reducing server costs. When distributing video screens waiting to be played to the user, the video screens waiting to be played can be retrieved from memory and sent to the video playback device at time intervals of 1 / fps. The video playback device does not need to perform any additional work and can simply decode and play the video.
[0085] In one embodiment, determining the video screen to be played from a first video screen stream and a second video screen stream based on the video playback parameters of a video playback device is: Based on the video playback parameters, determine whether the screen center point coordinates of the pending video screen are the same as the screen center point coordinates of the last video screen that was played, If they are different, the video stream in which the pending video screen is located is determined based on the video screen number of the pending video screen; if they are the same, the video stream in which the played historical video screen is located is determined based on the video stream in which the pending video screen is located. This includes selecting a video screen from the determined video screen stream as a video screen to be played, where the video screen number is equal to the video screen number of the video screen to be played.
[0086] In this embodiment, first, based on video playback parameters, it is determined whether the screen center point coordinates of the pending video screen are the same as the screen center point coordinates of the last video screen played. If they are the same, it is stated that the viewing angle will not be switched as the user's head moves; on the other hand, if they are different, it is stated that the viewing angle needs to be switched as the user's head moves. The purpose of this step is to determine whether it is necessary to select and send to the user a video screen with a different viewing angle from the last video screen played among the first and second video screen streams.
[0087] Here, it is necessary to briefly distinguish between the user's viewing angle and the viewing angle of the video screens in the first and second video screen streams. When a user watches on-demand video on a VR device, any head movement causes a change in the user's viewing angle, and consequently, the viewing angle displayed by the VR device also changes. However, due to the existence of the aforementioned field of view margin, after the user's viewing angle changes, the VR device does not necessarily select and send to the user a video screen with a different viewing angle from the last video screen played in the first and second video screen streams.
[0088] Therefore, here, we first determine, based on the video playback parameters, whether the viewing angle of the pending video screen is the same as the viewing angle of the last video screen played, that is, whether the coordinates of the center point of the pending video screen are the same as the coordinates of the center point of the last video screen played. If they are different, we indicate that we need to select a video screen with a different viewing angle from the last video screen played and send it to the user in the first video screen stream and the second video screen stream; on the other hand, if they are the same, we indicate that this is not necessary.
[0089] Next, if the viewing angles are different, the video stream in which the pending video screen is located is determined based on the video screen number of the pending video screen. If the viewing angles are the same, the video stream in which the pending video screen is located is determined based on the video stream in which the played history video screen is located. Finally, from the determined video streams, a video screen whose video screen number is equal to the video screen number of the pending video screen is selected and sent to the user as the pending video screen. The video screen number of the pending video screen can be obtained from the video playback parameters of the video playback device.
[0090] Thus, according to this embodiment, based on video playback parameters, it is possible to determine whether the coordinates of the center point of the video screen waiting to be played are the same as the coordinates of the center point of the last video screen that was played, and to select the video screen waiting to be played based on the determination result. This makes it possible to timely detect when it is necessary to switch the viewing angle and to deliver the necessary video screen to the user in a timely manner.
[0091] In one embodiment, determining whether the screen center point coordinates of a video screen waiting to be played are the same as the screen center point coordinates of the last video screen played, based on video playback parameters, Extracting the screen center point coordinates of the video playback device from the video playback parameters, The coordinates of the center point of the video playback device's screen and the field of view margin of the video screen relative to the field of view of the video playback device are used to determine the coordinates of the center point of the video screen waiting to be played back. This includes obtaining the screen center coordinates of the last video screen played and determining whether the screen center coordinates of the video screen waiting to be played are the same as the screen center coordinates of the last video screen played.
[0092] Here, the field of view margin and the coordinates of the center point of the last video screen played may be stored on the server or in the video playback parameters.
[0093] First, the screen center point coordinates of the video playback device are extracted from the video playback parameters. The screen center point coordinates of the video playback device refer to the pitch and yaw angle coordinates of the VR device, that is, the pitch and yaw angle coordinates of the user's viewpoint.
[0094] In one embodiment, determining the coordinates of the center point of the video screen waiting to be played based on the coordinates of the center point of the video playback device and the field of view margin of the video screen relative to the field of view of the video playback device is, specifically: The process involves calculating the screen center coordinates of the video screen waiting to be played, based on the screen center coordinates and field of view margins of the video playback device.
[0095] In one embodiment, the coordinates of the center point of the video screen waiting to be played are calculated based on the coordinates of the center point of the video playback device and the field of view margin using the following formula:
[0096]
number
[0097] In the above equation, pitch and yaw refer to the coordinates of the center point of the video playback device's screen, and Margin pitch and Margin yaw'pitch' and 'yaw' represent the field of view margins in the pitch direction and yaw direction, respectively, and refer to the coordinates of the center point of the video screen waiting to play.
[0098] In this embodiment, the coordinates of the center point of the last video screen played are also obtained, and this information may be stored in the server beforehand, or it may be stored in the video playback parameters. Next, it is determined whether the coordinates of the center point of the video screen waiting to be played are the same as the coordinates of the center point of the last video screen played. If they are the same, it is determined that the viewing angle of the video screen waiting to be played is the same as the viewing angle of the last video screen played. If they are different, it is determined that the viewing angle of the video screen waiting to be played is different from the viewing angle of the last video screen played.
[0099] Thus, according to this embodiment, the coordinates of the center point of the video screen waiting to be played can be determined based on the coordinates of the center point of the video playback device and the field of view margin of the video screen relative to the field of view of the video playback device. In other words, the viewing angle of the video screen waiting to be played can be accurately determined, thereby determining whether the viewing angle of the video screen waiting to be played is the same as the viewing angle of the last video screen that has been played, promptly detecting when it is necessary to switch the viewing angle, and ensuring that the necessary video screen is delivered to the user in a timely manner.
[0100] In one embodiment, when it is necessary to switch the viewing angle, the video screen stream in which the video screen waiting to be played is located is determined based on the video screen number of the video screen waiting to be played. Obtain a preset multiple of the length of the second screen group GOP of the second video screen stream, and calculate the remainder when the video screen number of the waiting video screen is divided by that preset multiple of the length, If the remainder is a pre-set value, a second video screen stream whose screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played is determined as the video screen stream where the video screen waiting to be played is located. This includes determining, if the remainder is not a pre-set value, that the video screen stream where the last played video screen is located is the video screen stream where the video screen awaiting playback is located.
[0101] Specifically, if the viewing angle of the video screen waiting to be played differs from the viewing angle of the last video screen played, a preset multiple of the length of the second screen group GOP of the second video screen stream is obtained. For example, if the second frame structure is an IP structure and one GOP contains one I-frame and one P-frame, the standard maximum number of frames in the frame group of the second frame structure is 2, meaning the length of the GOP is 2, and here multiples of 2 such as 2, 4, or 6 can be obtained.
[0102] Next, the remainder obtained by dividing the video screen number of the waiting video screen by its preset multiple is calculated. If the remainder is the set value, for example 1, then the second video screen stream whose screen center point coordinates are the same as the screen center point coordinates of the waiting video screen is determined as the video screen stream in which the waiting video screen is located. As can be seen from the above, each viewing angle has a first video screen stream and a second video screen stream, so the viewing angle of the waiting video screen also has a first video screen stream and a second video screen stream, and the angle is represented by the position coordinates of the screen center point. Here, when the remainder is the set value, the second video screen stream whose screen center point coordinates are the same as the screen center point coordinates of the waiting video screen, that is, the second video screen stream corresponding to the viewing angle of the waiting video screen, is determined as the video screen stream in which the waiting video screen is located.
[0103] Conversely, if the remainder is not the set value, the video screen stream where the last played video screen is located is determined to be the video screen stream where the pending video screen is located. The video screen stream where the last played video screen is located can be the first video screen stream corresponding to the last played video screen, or it can be the second video screen stream corresponding to the last played video screen.
[0104] Thus, in this embodiment, by calculating the remainder obtained by dividing the video screen number of the video screen waiting to be played by a preset multiple of the length of the second screen group GOP of the second video screen stream, and comparing the remainder with a set value, it is possible to quickly determine the video screen stream in which the video screen waiting to be played is located when a viewer angle switch is required. This process is simple, significantly improves the computational efficiency of the server, and reduces the computational pressure on the server.
[0105] In one embodiment, when there is no need to switch the viewing angle, determining the video screen stream in which the video screen waiting to be played is located based on the video screen stream in which the played historical video screen is located is: Obtain a preset multiple of the length of the second screen group GOP of the second video screen stream, and determine the video screen as the historical video screen if the interval between the video screen number of the historical playback and the video screen number of the waiting video screen is a preset multiple of that length. To determine whether the historical video screen is in a second video screen stream whose screen center coordinates are the same as those of the video screen waiting to be played, and whether it is in a first video screen stream whose screen center coordinates are the same as those of the video screen waiting to be played, This includes determining the video stream in which the video screen awaiting playback is located, based on the judgment result.
[0106] Specifically, we obtain a preset multiple of the length of the second screen group GOP of the second video screen stream. This process is the same as described above and will not be repeated.
[0107] Next, the video screen whose interval between the video screen number of the history playback and the video screen number of the video screen waiting to be played is a preset multiple of that length is determined to be the history video screen. For example, if the obtained value is 4 and the video screen number of the video screen waiting to be played is 16, then the history video screen will be the video screen with the history playback video screen number 12.
[0108] Next, it is determined whether the historical video screen is in a second video screen stream whose center point coordinates are the same as those of the video screen waiting to be played, and whether it is in a first video screen stream whose center point coordinates are the same as those of the video screen waiting to be played. This step is equivalent to determining whether the historical video screen is in a second video screen stream corresponding to the viewing angle of the video screen waiting to be played, and whether it is in a first video screen stream corresponding to the viewing angle of the video screen waiting to be played. Finally, based on the determination result, the video screen stream in which the video screen waiting to be played is located is determined.
[0109] Thus, according to this embodiment, by determining the historical video screen and analyzing the video screen stream in which the historical video screen is located, it is possible to quickly determine the video screen stream in which the video screen awaiting playback is located without the need to switch viewing angles. This process is simple and efficient, and reduces the computational pressure on the server.
[0110] In one embodiment, based on the above-described determination result, the video screen stream in which the video screen awaiting playback is located is determined as follows: If the historical video screen is in a second video screen stream where the screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played, or if the screen center point is in a first video screen stream where the screen center point is the same as the screen center point coordinates of the video screen waiting to be played, then the first video screen stream where the screen center point is the same as the screen center point coordinates of the video screen waiting to be played is determined as the video screen stream in which the video screen waiting to be played is located. The method includes determining the second video stream in which the center point coordinates are the same as the center point coordinates of the video screen waiting to be played, if the historical video screen is not in a second video stream in which the center point coordinates are the same as the center point coordinates of the video screen waiting to be played, and is not in a first video stream in which the center point coordinates are the same as the center point coordinates of the video screen waiting to be played, as the video stream in which the video screen waiting to be played is located.
[0111] Specifically, since each viewing angle has a first video screen stream and a second video screen stream, the viewing angle of the video screen waiting to be played also has a first video screen stream and a second video screen stream. Based on the judgment, if the historical video screen is in the second video screen stream where the screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played, or if the screen center point is in the first video screen stream where the screen center point is the same as the screen center point coordinates of the video screen waiting to be played, that is, if the historical video screen is in the second video screen stream corresponding to the viewing angle of the video screen waiting to be played, or in the first video screen stream corresponding to the viewing angle of the video screen waiting to be played, then the first video screen stream where the screen center point is the same as the screen center point coordinates of the video screen waiting to be played, i.e., the first video screen stream corresponding to the viewing angle of the video screen waiting to be played, is determined to be the video screen stream in which the video screen waiting to be played is located.
[0112] Conversely, if the historical video screen is not in a second video screen stream where the screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played, and is not in a first video screen stream where the screen center point is the same as the screen center point coordinates of the video screen waiting to be played, that is, if the historical video screen is not in a second video screen stream corresponding to the viewing angle of the video screen waiting to be played, and is not in a first video screen stream corresponding to the viewing angle of the video screen waiting to be played, then the second video screen stream where the screen center point is the same as the screen center point coordinates of the video screen waiting to be played, i.e., the second video screen stream corresponding to the viewing angle of the video screen waiting to be played, is determined as the video screen stream in which the video screen waiting to be played is located.
[0113] Thus, according to this embodiment, the video screen stream in which a video screen awaiting playback is located can be quickly determined based on the video screen stream in which the historical video screen is located. The process is simple and highly efficient, and the computational pressure on the server is low.
[0114] If there are multiple scaling ratio values for the video screen relative to a video frame, the initial video screen stream at the same screen center point position contains multiple substreams, each corresponding to a different scaling ratio value, and each substream is converted into a first video screen stream and a second video screen stream, respectively. In this embodiment of the disclosure, the scaling ratio of the video screen for each frame sent to the user's video playback device is usually the same, however, in the following cases, video screens with different scaling ratios may be sent to the user's video playback device: (1) In the case of a network carton, the user's video playback device can automatically select a scaling ratio value different from the currently used scaling ratio from among several pre-configured values, and the server sends the video screen to the video playback device based on that different value.
[0115] (2) When a user performs a scaling operation, the server determines an appropriate scaling ratio value in accordance with the operation and transmits the corresponding video screen.
[0116] (3) When the user pauses video playback, the server automatically switches to the minimum scaling ratio and increases the amount of information on the screen, so that the screen content can be changed even when the user pauses or moves their head in place.
[0117] (4) When the user moves their head quickly, if the server detects an acceleration that exceeds a preset threshold, it automatically switches to the minimum scaling ratio, thereby increasing the amount of screen information at the same field of view and reducing the probability of screen stretching.
[0118] (5) When playing the last frame of the video screen, if the user does not select loop playback, the server automatically switches to the minimum scaling ratio to increase the amount of image information the user can see.
[0119] In this embodiment, considering that scaling may be switched as described above, the process of the above method determines, based on the video playback parameters, whether the scaling corresponding to the video screen being played is the same as the scaling of the last video screen played, before determining whether the screen center point coordinates of the waiting video screen are the same as the screen center point coordinates of the last video screen played. If they are the same, Based on the aforementioned video playback parameters, it is determined whether the video center point coordinates of the video waiting to be played are the same as the screen center point coordinates of the last video screen that was played. If they are different, the video screen stream in which the pending video screen is located is determined based on the video screen number of the pending video screen; however, if they are the same, the video screen stream in which the played historical video screen is located is determined based on the video screen stream in which the pending video screen is located. From the determined video screen stream, a video screen whose video screen number is equal to the video screen number of the waiting video screen is selected as the waiting video screen.
[0120] If different, Based on the video screen number of the video screen waiting to be played, the video screen stream in which the video screen waiting to be played is located is determined, and from the determined video screen stream, a video screen whose video screen number is equal to the video screen number of the video screen waiting to be played is selected as the video screen waiting to be played.
[0121] See the above explanation for details of this process. It will not be repeated here.
[0122] Figure 7 is a schematic flowchart of a flowchart for selecting a video screen to play, provided by one embodiment of the present disclosure. As shown in Figure 7, this process includes the following steps:
[0123] In step S702, the user's video playback parameters are obtained, and the video screen number of the video waiting to be played is extracted from the video playback parameters.
[0124] In step S704, it is determined whether the video screen number of the video waiting to be played is 3 or greater.
[0125] If the value is 3 or greater, step S706 is performed; otherwise, step S718 is performed.
[0126] In step S706, based on the video playback parameters, it is determined whether the scaling ratio is the same between the viewing angle of the video screen waiting to be played and the viewing angle of the last video screen that has been played.
[0127] If they are the same, step S708 is performed; otherwise, step S712 is performed.
[0128] In step S708, the aforementioned historical video screen is determined, and it is determined whether the historical video screen is in a second video stream that corresponds to the viewing angle of the video screen waiting to be played.
[0129] If in the second video stream, perform step S718; otherwise, perform step S710.
[0130] In step S710, it is determined whether the historical video screen is in the first video screen stream that corresponds to the viewing angle of the video screen waiting to be played.
[0131] If the first video screen stream is in progress, step S718 is performed; otherwise, step S716 is performed.
[0132] In step S712, the preset multiple of the length of the second screen group GOP of the second video screen stream is obtained, and the remainder obtained by dividing the video screen number of the video screen waiting to be played by that preset multiple of its length is calculated to determine whether it matches the set value.
[0133] If the remainder is the set value, step S716 is executed; otherwise, step S714 is executed.
[0134] In step S714, the video screen stream containing the last video screen that was played is determined to be the video screen stream containing the video screen waiting to be played.
[0135] In step S716, a second video screen stream corresponding to the viewing angle of the video screen waiting to be played is determined as the video screen stream in which the video screen waiting to be played is located.
[0136] In step S718, a first video screen stream corresponding to the viewing angle of the video screen waiting to be played is determined as the video screen stream in which the video screen waiting to be played is located.
[0137] In step S720, from the determined video screen stream, a video screen whose video screen number is equal to the video screen number of the video screen waiting to be played is selected as the video screen waiting to be played.
[0138] In Figure 7, if the video screen number of the video waiting to play is less than 3, the on-demand video has just started playing, so the user is not allowed to switch the viewing angle, and normally the user would not switch the viewing angle at this point. Therefore, step S718 is executed, and the first video screen stream corresponding to the viewing angle of the video waiting to play is determined as the video screen stream in which the video waiting to play is located. The specific process in Figure 7 can be found in the previous explanation and will not be repeated here.
[0139] The first frame structure of the first video screen stream is the default video playback frame structure of the video playback device, and the second frame structure of the second video screen stream is the auxiliary video playback frame structure of the video playback device. When sending video screens with different viewing angles to the user, or when it is not possible to play a video screen in the first video screen stream, the video playback device will preferentially select a video screen from the second video screen stream, and ultimately use the video screen in the first video screen stream to play the video.
[0140] Figure 8 is a schematic diagram of a user-switched video screen provided by one embodiment of the present disclosure, and Figure 9 is a schematic diagram of a scene in which a video screen waiting to be played is selected, provided by one embodiment of the present disclosure. In Figures 8 and 9, S represents the viewing angle and F represents the video screen number. In Figure 9, the IPP flow is the first video screen stream and the IP flow is the second video screen stream. Referring to Figure 8, when the user attempts to play the video screen of the third frame, i.e., when the video screen waiting to be played is the video screen of the third frame, the user moves their head over a wide range, thereby causing the server to perform a viewing angle switching operation (also called a streaming operation). Similarly, when the user attempts to play the video screen of the sixth frame, i.e., when the video screen waiting to be played is the video screen of the sixth frame, the user moves their head over a wide range, causing the server to perform a viewing angle switching operation.
[0141] Referring to Figures 9 and 7, when switching from the video screen of the second frame to the video screen of the third frame, the video screen of the second frame is in the first video screen stream. Following the method process in Figure 7, it is necessary to perform step S712, and when the length of the second frame structure is 2, it is determined that the remainder when video screen number 3, which is the video screen to obtain the waiting video screen, is divided by the length 2 is the set value 1, and steps S716 and S720 are performed to select the video screen of the third frame from the second video screen stream corresponding to the viewing angle of the waiting video screen (i.e., viewing angle 2) as the waiting video screen.
[0142] Referring to Figure 9, the video screen of the third frame from the second video screen stream corresponding to viewing angle 2 is selected as the video screen to be played and played. Since the video screen to be played is an I-frame, the P-frame following the I-frame is then played. After the playback of the P-frame is finished, the program jumps to the first video screen stream corresponding to viewing angle 2 and continues playback.
[0143] Referring to Figures 9 and 7, when switching from the video screen of the fifth frame to the video screen of the sixth frame, the video screen of the fifth frame is in the first video screen stream. Following the method process in Figure 7, step S712 must be executed. If the length of the frame group of the pre-set second frame structure is 2, it is determined that the remainder when video screen number 6, which is the video screen to obtain for playback, is divided by 2 is not the set value of 1. Steps S714 and S720 are then executed, and the video screen of the sixth frame is selected as the video screen for playback from the video screen stream where the last played video screen is located (i.e., the first video screen stream corresponding to viewing angle 2).
[0144] Referring to Figure 9, immediately after the sixth frame finishes playing, playback continues by switching to the video screen of the seventh frame in the second video screen stream corresponding to viewing angle 3.
[0145] In Figure 9, each time the next frame of the video screen is played, the next frame of the video screen can be selected by referring to the method process in Figure 7. For example, when playing the fourth frame of the video screen, following the method process in Figure 7, the historical video screen is the second frame of the video screen and is not located in the first and second video screen streams corresponding to the viewing angle of the video screen to be played. Therefore, steps S716 and S720 are performed to select the video screen to be played from the second video screen stream corresponding to the viewing angle of the video screen to be played, thus selecting the fourth frame of the video screen in the drawing.
[0146] As can be seen in Figure 9, when switching from the second video screen stream to the first video screen stream within the same viewing angle, it is possible to switch from a P-frame of the second video screen stream to an I-frame or P-frame of the first video screen stream. Within different viewing angles, it is possible to switch from a P-frame of the first video screen stream in the previous viewing angle to an I-frame of the second video screen stream in the next viewing angle.
[0147] As can be seen from the above, when playing on-demand video using the second video screen stream, playback can be directly switched from the P-frame of the second video screen stream to the P-frame or I-frame of the first video screen stream with the same viewing angle. When viewing angle switching occurs in this way, the delay is small. For example, in the case of the second viewing angle switch in Figure 8, the viewing angle switch is finally achieved in the seventh frame, resulting in a delay of one frame. This significantly reduces the delay of viewing angle switching compared to when the auxiliary code stream consists entirely of I-frames.
[0148] Based on the embodiments described herein: 1. By generating a VAM screen that is slightly larger than the user's viewport but much smaller than the 8K panoramic screen, the bandwidth required for the server to send panoramic video to the VR device during on-demand panoramic video playback on the VR device can be reduced, improving the user's on-demand panoramic video experience on the VR device, and the video screen update delay time can be controlled to within 25 ms.
[0149] 2. By pre-generating and storing the first and second video screen streams, when a video is requested, only the video screen needs to be delivered. This significantly increases the amount of concurrent user activity, reduces the server's computing power requirements and operating costs, and further reduces costs by using a server based on the ARM architecture.
[0150] Figure 10 is a schematic diagram of the configuration of a video processing device provided by one embodiment of the present disclosure. As shown in Figure 10, this device is A video acquisition unit 1001 acquires a video stream awaiting processing, which includes multiple video frames having a field of view larger than the field of view of the video playback device. A video cutting unit 1002 cuts each of the plurality of video frames into a plurality of video screens so that the field of view of the video screens is larger than the field of view of the video playback device, For video screens among the plurality of video screens whose center point is in the same position, a first video screen stream and a second video screen stream are generated in the time order of the playback timestamps of the video screens, the first screen group GOP of the first video screen stream is a standard GOP, and the length of the second screen group GOP of the second video screen stream is smaller than the length of the first screen group GOP of the first video screen stream, the video generation unit 1003, The system includes a video distribution unit 1004 that determines a video screen awaiting playback from the first video screen stream and the second video screen stream based on the video playback parameters of the video playback device.
[0151] As an option, the video cutting unit 1002 specifically: Obtaining video screen cut parameters including the resolution of the video screen, the scaling ratio of the video screen to the video frame, the margin of the field of view of the video screen relative to the field of view of the video playback device, and the field of view interval between the field of view of the video playback device and the adjacent video screen, This is applied to cutting each of the multiple video frames into multiple video screens based on the aforementioned video screen cut parameters.
[0152] Optionally, it further includes a first parameter determination unit, The acquisition of a first range of values for the resolution of the video screen, a second range of values for the scaling ratio, and a third range of values for the field of view margin, A constraint used to express that the ratio result of the resolution of the video screen and the scaling ratio is equal to the field of view margin by subtracting the field of view of the video playback device, wherein constraints on the resolution of the video screen, the scaling ratio, and the field of view margin are obtained. This applies to determining the resolution of the video screen that satisfies the constraints, the scaling ratio that satisfies the constraints, and the field of view margin that satisfies the constraints, within the range of the first value, the range of the second value, and the range of the third value, respectively.
[0153] Optionally, it further includes a second parameter determination unit. Based on the aforementioned field of view margin and the field of view of each video frame, the viewing angle is applied to determine the viewing angle interval between adjacent video screens.
[0154] As an option, the second parameter determination unit described above specifically includes: A first constraint condition is obtained for the aforementioned viewing angle interval, the first constraint condition being that the viewing angle interval is less than or equal to the aforementioned field of view margin, Obtain a second constraint condition on the viewing angle interval, which indicates that the field of view of each video frame is divisible by the viewing angle interval. This applies to determining the visual angle interval based on the first constraint condition of the visual angle interval and the second constraint condition of the visual angle interval.
[0155] As an option, the video cutting unit 1002 is more specifically: The aspect ratio of the video screen is determined based on the aspect ratio of the video playback device and the aspect ratio margin. For each of the plurality of video frames, the video frame is mapped to a unit sphere, and the mapped image is scaled based on the scaling ratio. This method applies to each of the plurality of video frames, cutting the scaled image of the video frame based on the field of view, resolution, and viewing angle interval of the video frame, in order to obtain the video screen at each viewing angle of the video frame.
[0156] As an option, the video cutting unit 1002 is more specifically: The center point of the video screen is aligned with the center point of the image after scaling of the video frame, the image after scaling of the video frame is cropped according to the resolution and the field of view of the video screen, and the video screen is obtained at the reference viewing angle of the video frame. This method applies to moving the screen center point of the video screen onto the scaled image of the video frame based on the aforementioned viewing angle interval, cutting the scaled image of the video frame according to the resolution and the field of view of the moved video screen, and obtaining the video screen at other viewing angles of the video frame.
[0157] As an option, the video generation unit 1003 specifically: For the video screens among the aforementioned multiple video screens that have the same center point position, the initial video screen stream is obtained by arranging the video screens in the time order of their playback timestamps. The initial video screen stream is converted into a first video screen stream and a second video stream, respectively, wherein the first screen group GOP of the first video screen stream includes one I-frame and at least two P-frames, and the second screen group GOP of the second video screen stream includes one I-frame and at most two P-frames.
[0158] As an option, it includes additional memory units. For the first and second video screen streams generated by video screens with the same center point position among the multiple video screens, a name is assigned to each video screen in the first video screen stream and each video screen in the second video screen stream according to the center point position, the identifier of the first screen group GOP, the identifier of the second screen group GOP, and the video screen number setting rule, and name information is obtained. This applies to storing the first and second video screen streams based on the naming information of each video screen in the first video screen stream and the naming information of each video screen in the second video screen stream.
[0159] As an option, the video distribution unit 1004 specifically: Based on the aforementioned video playback parameters, it is determined whether the coordinates of the center point of the video screen waiting to be played are the same as the coordinates of the center point of the last video screen that was played. If they are different, the video stream in which the pending video screen is located is determined based on the video screen number of the pending video screen; if they are the same, the video stream in which the played historical video screen is located is determined based on the video stream in which the pending video screen is located. This applies to selecting a video screen from the determined video screen stream as the video screen to be played, where the video screen number is equal to the video screen number of the video screen to be played.
[0160] As an option, the video distribution unit 1004 is, more specifically: Extracting the coordinates of the screen center point of the video playback device from the aforementioned video playback parameters, The coordinates of the center point of the video playback device's screen and the field of view margin of the video screen relative to the field of view of the video playback device's field of view are used to determine the coordinates of the center point of the video screen waiting to be played back. This is applied to obtaining the screen center point coordinates of the last video screen that was played and determining whether the screen center point coordinates of the video screen waiting to be played are the same as the screen center point coordinates of the last video screen that was played.
[0161] As an option, the video distribution unit 1004 is, more specifically: Obtain a preset multiple of the length of the second screen group GOP of the second video screen stream, calculate the remainder when the video screen number of the video screen waiting to be played is divided by the preset multiple of the length, If the remainder is a set value, the second video screen stream whose screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played is determined to be the video screen stream in which the video screen waiting to be played is located. If the remainder is not a set value, the video screen stream in which the last played video screen is located is determined to be the video screen stream in which the video screen awaiting playback is located.
[0162] As an option, the video distribution unit 1004 is, more specifically: Obtain a preset multiple of the length of the second screen group GOP of the second video screen stream, and determine the video screen whose interval between the video screen number of the historical playback and the video screen number of the waiting video screen is a preset multiple of the length as the historical video screen, The process involves determining whether the historical video screen is in the second video screen stream whose screen center point coordinates are the same as those of the video screen waiting to be played, and whether it is in the first video screen stream whose screen center point coordinates are the same as those of the video screen waiting to be played. This applies to determining the video screen stream in which the video screen awaiting playback is located, based on the judgment result.
[0163] As an option, the video distribution unit 1004 is, more specifically: If the historical video screen is in the second video screen stream whose screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played, or if the screen center point is in the first video screen stream whose screen center point is the same as the screen center point coordinates of the video screen waiting to be played, then the first video screen stream whose screen center point is the same as the screen center point coordinates of the video screen waiting to be played is determined to be the video screen stream in which the video screen waiting to be played is located. If the historical video screen is not in the second video screen stream whose screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played, and is not in the first video screen stream whose screen center point is the same as the screen center point coordinates of the video screen waiting to be played, then the second video screen stream whose screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played is determined to be the video screen stream in which the video screen waiting to be played is located.
[0164] Optionally, each of the aforementioned video frames is a panoramic video frame.
[0165] The video processing apparatus in the embodiments of this disclosure can implement various processes of the above-described embodiments of the video processing method, and can achieve the same effects and functions, so it will not be repeated here.
[0166] One embodiment of the present disclosure also provides an electronic device. Figure 11 is a schematic diagram of the configuration of an electronic device provided by one embodiment of the present disclosure. As shown in Figure 11, the electronic device can vary relatively significantly in terms of arrangement or performance. It may include one or more processors 1101 and a memory 1102, the memory 1102 may store one or more application programs or data. Here, the memory 1102 may be temporary or persistent storage. The application program stored in the memory 1102 may include one or more modules (not shown) that can contain a set of computer executable instructions within the electronic device. Furthermore, the processor 1101 can communicate with the memory 1102 and be configured in the electronic device to execute a set of computer executable instructions in the memory 1102. The electronic device may also include one or more power supplies 1103, one or more wired or wireless network interfaces 1104, one or more input or output interfaces 1105, one or more keyboards 1106, and the like.
[0167] In one specific embodiment, the electronic device is a video processing unit, specifically a background server, and includes a processor and memory arranged to store computer executable instructions, which, when executed, cause the processor to perform the following processes: Obtain a video stream awaiting processing, which contains multiple video frames with a field of view larger than the field of view of the video playback device.
[0168] Each of the aforementioned multiple video frames is cut into multiple video screens, and the aspect ratio of the video screens is larger than the aspect ratio of the video playback device.
[0169] For video screens among the plurality of video screens that have the same center point position, a first video screen stream and a second video screen stream are generated in the time order of the playback timestamps of the video screens, the first screen group GOP of the first video screen stream is a standard GOP, and the length of the second screen group GOP of the second video screen stream is smaller than the length of the first screen group GOP of the first video screen stream.
[0170] Based on the video playback parameters of the video playback device, the video screen to be played is determined from the first video screen stream and the second video screen stream.
[0171] The video processing apparatus in the embodiments of this disclosure can implement various processes of the above-described embodiments of the video processing method, and can achieve the same effects and functions, so it will not be repeated here.
[0172] Another embodiment of the present disclosure further provides a computer-readable storage medium for storing computer-executable instructions that, when executed by a processor, accomplish the following processes:
[0173] Obtain a video stream awaiting processing, which contains multiple video frames with a field of view larger than the field of view of the video playback device.
[0174] Each of the aforementioned multiple video frames is cut into multiple video screens, and the aspect ratio of the video screens is larger than the aspect ratio of the video playback device.
[0175] For video screens among the plurality of video screens that have the same center point position, a first video screen stream and a second video screen stream are generated in the time order of the playback timestamps of the video screens, the first screen group GOP of the first video screen stream is a standard GOP, and the length of the second screen group GOP of the second video screen stream is smaller than the length of the first screen group GOP of the first video screen stream.
[0176] Based on the video playback parameters of the video playback device, the video screen to be played is determined from the first video screen stream and the second video screen stream.
[0177] The storage medium in the embodiments of this disclosure can implement various processes of the video processing method embodiments described above, and can achieve the same effects and functions, so it will not be repeated here.
[0178] In various embodiments of this disclosure, the computer-readable storage medium includes read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0179] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, and switches) or software improvements (improvements to method processes). However, with technological advancements, many current method process improvements can now be considered direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method process into the hardware circuit. Therefore, it cannot be said that improvements to a method process cannot be realized with hardware entity modules. For example, programmable logic devices (PLDs) (e.g., field programmable gate arrays (FPGAs)) are such integrated circuits, and their logic function is determined by the user programming the device. There is no need to have a chip manufacturer design and manufacture a dedicated integrated circuit chip; designers can program and "integrate" a single digital system into a PLD. Furthermore, instead of manually creating integrated circuit chips, these programs are now often implemented using "logic compiler" software, which is similar to the software compilers used for program development, and the original code before compilation must also be written in a specific programming language.These are called Hardware Description Languages (HDLs), and there isn't just one type of HDL. There are several types, including ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. As those skilled in the art will know, by simply programming a method process with a little logic using one of the above hardware description languages and then programming it into an integrated circuit, it is easy to obtain the hardware circuit that realizes that logic method process.
[0180] A controller can be implemented in any suitable manner, for example, in the form of a microprocessor or processor, a computer-readable medium for storing computer-readable program code (e.g., software or firmware) executed by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, or embedded microcontrollers. Examples of controllers include, but are not limited to, microcontrollers such as the ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. A memory controller can be implemented as part of the control logic for memory. As those skilled in the art will see, in addition to implementing a controller in a purely computer-readable program code manner, it is entirely possible to implement the same function of a controller in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logic programming method steps. Thus, such a controller can be considered a hardware component, and the devices contained within it for implementing various functions can also be considered components within the hardware component. Alternatively, the device for implementing various functions may be a software module for realizing the method, or it may be a configuration within a hardware component.
[0181] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products having some function. One typical implement is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a mobile phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0182] For the sake of explanation, when describing the above device, we will divide it into various units according to their function and describe each unit separately. Of course, when implementing the embodiments of this disclosure, the functions of each unit can be realized with the same or multiple software and / or hardware.
[0183] As those skilled in the art will see, one or more embodiments of the present disclosure can provide a method, system, or computer program product. Accordingly, one or more embodiments of the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of the present disclosure may take the form of a computer program product that runs on one or more computer-available storage media (including, but not limited to, disk memory, CD-ROM, optical memory, etc.) which contain computer-available program code.
[0184] This disclosure will be described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. Computer program instructions can implement each process and / or block in the flowcharts and / or block diagrams, and combinations of processes and / or blocks in the flowcharts and / or block diagrams. These computer program instructions can be provided to a processor of a general-purpose computer, a dedicated computer, an embedded processor, or other programmable data processing device, so as to build a machine, which causes instructions executed by the computer or other programmable data processing device processor to produce apparatus for implementing one or more processes in the flowchart and / or one or more blocks in the block diagram.
[0185] These computer program instructions can also be stored in computer-readable memory, which can guide a computer or other programmable data processing device to operate in a particular manner, thereby generating a product that includes instruction devices that implement the functions specified in one or more processes in a flowchart and / or one or more blocks in a block diagram.
[0186] These computer program instructions may also be loaded into a computer or other programmable data processing device, thereby enabling the execution of a series of operational steps on the computer or other programmable device to generate computer-implemented processing, the instructions executed on the computer or other programmable device providing steps to implement one or more processes in a flowchart and / or one or more blocks in a block diagram.
[0187] The terms “include” (or “encompass” in Chinese), “include” (or “contain” in Chinese), or any other variations are intended to cover the non-exclusive “include.” This means that a process, method, product, or device containing a set of elements includes not only those elements, but also other elements not explicitly listed, or elements specific to such a process, method, product, or device. Unless otherwise specified, an element limited by the phrase “…includes one” does not preclude other identical elements from the process, method, product, or device containing the aforementioned element.
[0188] One or more embodiments of the present disclosure can be described in the general context of computer executable instructions executed by a computer, such as program modules. Generally, a program module includes routines, programs, objects, components, data structures, etc., that perform a particular task or realize a particular abstract data type. One or more embodiments of the present disclosure can also be implemented in a distributed computing environment in which tasks are performed by remote processing units connected via a communication network. In a distributed computing environment, program modules can reside on local and remote computer storage media, including storage devices.
[0189] Each embodiment in this disclosure is described in a stepwise manner, with similar parts between embodiments being able to refer to one another, and each embodiment focusing on parts that differ from the others. In particular, the system embodiments are almost identical to the method embodiments and are therefore relatively simple to describe; relevant parts should be referred to the description of the method embodiments.
[0190] The foregoing are merely examples of the present disclosure and are not intended to limit the present disclosure. The present disclosure is subject to various modifications and changes for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present disclosure should be included within the scope of the claims of the present disclosure.
Claims
1. The process involves obtaining a video stream awaiting processing that includes multiple video frames, wherein the aspect ratio of each of the multiple video frames is greater than the aspect ratio of the video playback device. Cutting each of the aforementioned plurality of video frames into multiple video screens, wherein the field of view of the video screens is larger than the field of view of the video playback device, For video screens among the plurality of video screens where the position of the center point of the screen is the same, a first video screen stream and a second video screen stream are generated in the time order of the playback timestamps of the video screens, wherein the first screen group GOP of the first video screen stream is a standard GOP, and the length of the second screen group GOP of the second video screen stream is smaller than the length of the first screen group GOP of the first video screen stream. A video processing method characterized by including determining a video screen to be played from the first video screen stream and the second video screen stream based on the video playback parameters of the video playback device.
2. Cutting each of the aforementioned multiple video frames into multiple video screens is, Obtaining video screen cut parameters including the resolution of the video screen, the scaling ratio of the video screen to the video frame, the field of view margin of the video screen to the field of view of the video playback device, and the field of view interval between the field of view of the video playback device and adjacent video screens, The method according to claim 1, characterized by comprising cutting each of the plurality of video frames into a plurality of video screens based on the video screen cut parameter.
3. The resolution of the video screen, the scaling ratio of the video screen to the video frame, and the margin of the field of view of the video screen to the field of view of the video playback device are: The steps include obtaining a range of a first value for the resolution of the video screen, a range of a second value for the scaling ratio, and a range of a third value for the field of view margin, A constraint used to express that the ratio result of the resolution of the video screen and the scaling ratio is equal to the field of view margin, wherein the constraint is obtained by subtracting the field of view of the video playback device from the ratio result of the resolution of the video screen and the scaling ratio, and the constraint is obtained for the resolution of the video screen, the scaling ratio and the field of view margin. The method according to claim 2, characterized in that the resolution of the video screen that satisfies the constraints, the scaling ratio that satisfies the constraints, and the field of view margin that satisfies the constraints are determined by the step of determining within the range of a first value, a range of a second value, and a range of a third value, respectively.
4. The aforementioned visual angle interval is, The method according to claim 2, characterized in that the viewing angle is determined by the step of determining the viewing angle interval between adjacent video screens based on the viewing angle margin and the viewing angle of each video frame.
5. Determining the viewing angle interval between adjacent video screens based on the aforementioned viewing angle margin and the viewing angle of each video frame means that, Obtaining a first constraint condition for the aforementioned viewing angle interval, which indicates that the viewing angle interval is less than or equal to the field of view margin, A second constraint condition is obtained that the field of view of each video frame is divisible by the aforementioned field of view interval, The method according to claim 4, characterized in that it includes determining the viewing angle interval based on a first constraint condition for the viewing angle interval and a second constraint condition for the viewing angle interval.
6. Cutting each video frame in the plurality of video frames into a plurality of video screens based on the aforementioned video screen cut parameters is: The aspect ratio of the video screen is determined based on the aspect ratio of the video playback device and the aspect ratio margin. For each of the plurality of video frames, the video frame is mapped to a unit sphere, and the mapped image is scaled based on the scaling ratio. The method according to any one of claims 2 to 5, characterized in that, for each of the plurality of video frames, the scaled image of the video frame is cut based on the field of view of the video screen, the resolution, and the viewing angle interval, to obtain the video screen at each viewing angle of the video frame.
7. Based on the field of view, resolution, and viewing angle interval of the video screen, cropping the image after scaling the video frame to obtain the video screen at each viewing angle of the video frame is: The center point of the video screen is aligned with the center point of the image after scaling of the video frame, the image after scaling of the video frame is cropped according to the resolution and the field of view of the video screen, and the video screen is obtained at the reference viewing angle of the video frame. The method according to claim 6, characterized by comprising: moving the screen center point of the video screen onto the scaled image of the video frame based on the viewing angle interval; cutting the scaled image of the video frame according to the resolution and the viewing angle of the moved video screen to obtain the video screen at other viewing angles of the video frame.
8. For video screens among the aforementioned plurality of video screens where the position of the screen center point is the same, generating a first video screen stream and a second video screen stream in the time order of the playback timestamps of those video screens is: For video screens among the aforementioned multiple video screens where the position of the center point of the screen is the same, the initial video screen stream is obtained by arranging the video screens in the time order of the playback timestamps of those video screens. The method according to any one of claims 1 to 5, characterized in that the initial video screen stream is converted into a first video screen stream and a second video stream, the first screen group GOP of the first video screen stream includes one I-frame and at least two P-frames, and the second screen group GOP of the second video screen stream includes one I-frame and at most two P-frames.
9. After generating the first video screen stream and the second video screen stream, the method further: For the first and second video screen streams generated by video screens with the same center point position among the plurality of video screens, a name is assigned to each video screen in the first video screen stream and each video screen in the second video screen stream according to the center point position, the identifier of the first screen group GOP, the identifier of the second screen group GOP, and the video screen number setting rule, and name information is obtained. The method according to any one of claims 1 to 5, characterized in that it includes storing the first and second video screen streams based on the naming information of each video screen in the first video screen stream and the naming information of each video screen in the second video screen stream.
10. Determining a video screen to be played from the first video screen stream and the second video screen stream based on the video playback parameters of the video playback device is: Based on the aforementioned video playback parameters, it is determined whether the coordinates of the center point of the video screen waiting to be played are the same as the coordinates of the center point of the last video screen that was played. If they are different, the video stream in which the pending video screen is located is determined based on the video screen number of the pending video screen; if they are the same, the video stream in which the played historical video screen is located is determined based on the video stream in which the pending video screen is located. The method according to any one of claims 1 to 5, characterized in that it includes selecting a video screen from a determined video screen stream as the video screen to be played, the video screen number being the same as the video screen number of the video screen to be played.
11. Based on the aforementioned video playback parameters, determining whether the screen center point coordinates of the video screen waiting to be played are the same as the screen center point coordinates of the last video screen that was played is: Extracting the coordinates of the screen center point of the video playback device from the aforementioned video playback parameters, The coordinates of the center point of the video playback device's screen and the field of view margin of the video screen relative to the field of view of the video playback device's field of view are used to determine the coordinates of the center point of the video screen waiting to be played back. The method according to 10, characterized by comprising obtaining the screen center point coordinates of the last video screen that has been played, and determining whether the screen center point coordinates of the video screen waiting to be played are the same as the screen center point coordinates of the last video screen that has been played.
12. Determining the video stream in which the pending video screen is located based on the video screen number of the pending video screen means that Obtain a preset multiple of the length of the second screen group GOP of the second video screen stream, and calculate the remainder when the video screen number of the video screen waiting to be played is divided by the preset multiple of the length. If the remainder is a set value, the second video screen stream whose screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played is determined to be the video screen stream in which the video screen waiting to be played is located. The method according to 10, characterized in that, if the remainder is not a set value, the video screen stream in which the last played video screen is located is determined to be the video screen stream in which the video screen awaiting playback is located.
13. Determining the video stream in which the video screen waiting to be played is located, based on the video screen stream in which the previously played historical video screen is located, The preset multiple of the length of the second screen group GOP of the second video screen stream is obtained, and the video screen whose interval between the video screen number of the historical playback and the video screen number of the waiting video screen is a preset multiple of the length is determined to be the historical video screen, The process involves determining whether the historical video screen is in the second video screen stream whose screen center point coordinates are the same as those of the video screen waiting to be played, and whether it is in the first video screen stream whose screen center point coordinates are the same as those of the video screen waiting to be played. The method according to 10, characterized in that it includes determining the video screen stream in which the video screen waiting to be played is located, based on the determination result.
14. Based on the aforementioned determination result, determining the video screen stream in which the video screen waiting to be played is: If the historical video screen is in the second video screen stream whose screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played, or if the screen center point is in the first video screen stream whose screen center point is the same as the screen center point coordinates of the video screen waiting to be played, then the first video screen stream whose screen center point is the same as the screen center point coordinates of the video screen waiting to be played is determined to be the video screen stream in which the video screen waiting to be played is located. The method according to 13, characterized in that, if the historical video screen is not in the second video screen stream in which the screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played, and the screen center point is not in the first video screen stream in which the screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played, then the second video screen stream in which the screen center point coordinates are the same as the screen center point coordinates of the video screen waiting to be played is determined to be the video screen stream in which the video screen waiting to be played is located.
15. The method according to any one of claims 1 to 14, characterized in that each of the plurality of video frames is a panoramic video frame.
16. A video acquisition unit for acquiring a video stream awaiting processing, which includes multiple video frames, wherein the aspect ratio of each video frame among the multiple video frames is greater than the aspect ratio of a video playback device, A video cutting unit for cutting each of the plurality of video frames into a plurality of video screens, wherein the field of view of the video screen is larger than the field of view of the video playback device, A video generation unit for generating a first video screen stream and a second video screen stream in the time order of the playback timestamps of multiple video screens that have the same screen center point position, wherein the first screen group GOP of the first video screen stream is a standard GOP, and the length of the second screen group GOP of the second video screen stream is smaller than the length of the first screen group GOP of the first video screen stream, A video processing device comprising: a video distribution unit for determining a video screen awaiting playback from the first video screen stream and the second video screen stream based on the video playback parameters of the video playback device.
17. An electronic device comprising a processor and a memory arranged to store computer executable instructions, wherein when the computer executable instructions are executed, the memory causes the processor to perform the steps of the method according to any one of claims 1 to 15.
18. A computer-readable storage medium characterized by storing computer-executable instructions that, when executed by a processor, realize a step of the method described in any one of claims 1 to 15.
Citation Information
Patent Citations
Bitstream structure for viewport-based streaming with a fallback bitstream
US20210377527A1
Method, apparatus and computer program product providing for extended margins around a viewport for immersive content
WO2021073940A1