A method, device, medium and equipment for pushing live panoramic video
By setting picture quality-sensitive tags on the prediction client and combining the non-predictive client historical trajectory of the delay-sensitive tag for perspective prediction and encoding, the problem that cannot meet the video quality requirements of different clients in the prior art is solved, and higher video push streaming accuracy is achieved.
Patent Information
- Application Number
- CN202211329369.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-10-27
AI Technical Summary
In the process of video streaming, the prior art cannot perform differential streaming for clients with different video playback needs, especially the needs of clients with high video quality requirements.
By setting picture quality sensitive tags on the prediction client, obtaining its historical viewing trajectory, and combining the historical viewing trajectory of the non-predictive client with the delay sensitive tag, perspective prediction and panoramic video encoding are performed to improve the accuracy of video push streaming.
It improves the accuracy of video encoding and streaming, and meets the needs of clients with high video quality requirements.
Smart Images

Figure CN115665503B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to video processing technologies, and in particular, to a method, apparatus, medium, and device for pushing live panoramic videos Background Art
[0002] During the process of video pushing, for clients with different video playback requirements, the same pushing strategy is adopted for video pushing.
[0003] The above pushing strategy cannot perform differential pushing for clients with different requirements. Especially for clients with high requirements for video image quality, the above pushing method cannot meet the user's requirements for video image quality. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, medium, and device for pushing live panoramic videos to improve the accuracy of video pushing.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for pushing a live panoramic video, including:
[0006] When a predicted client is set with a video quality sensitive label, obtaining a first historical viewing trajectory of the predicted client, and obtaining second historical viewing trajectories of multiple non-predicted clients set with a latency sensitive label, where the playback latency of the non-predicted clients is less than the playback latency of the predicted client;
[0007] Based on the first historical viewing trajectory and multiple second historical viewing trajectories, obtaining a predicted viewing perspective of the predicted client;
[0008] Performing panoramic video encoding based on the viewing perspective, and pushing the encoded video data to the predicted client.
[0009] In a second aspect, an embodiment of the present disclosure further provides a device for pushing a live panoramic video, including:
[0010] A data acquisition module, configured to obtain a first historical viewing trajectory of the predicted client when the predicted client is set with a video quality sensitive label, and obtain second historical viewing trajectories of multiple non-predicted clients set with a latency sensitive label, where the playback latency of the non-predicted clients is less than the playback latency of the predicted client;
[0011] A viewing perspective prediction module, configured to obtain a predicted viewing perspective of the predicted client based on the first historical viewing trajectory and multiple second historical viewing trajectories;
[0012] A video encoding module, configured to perform panoramic video encoding based on the viewing perspective;
[0013] A video streaming module, configured to stream the encoded video data to the prediction client.
[0014] In a third aspect, an embodiment of the present disclosure further provides an electronic device, which includes:
[0015] One or more processors;
[0016] A storage device, configured to store one or more programs,
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for streaming a live panoramic video provided by the present invention.
[0018] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the method for streaming a live panoramic video provided by the present invention when executed by a computer processor.
[0019] In the technical solution of the embodiment of the present disclosure, for a prediction client set with a picture quality sensitive tag, during the process of streaming a live panoramic video, the first historical viewing trajectory of the prediction client and the second historical viewing trajectories of multiple non-prediction clients set with multiple delay sensitive tags are obtained. By obtaining the second historical viewing trajectories of multiple non-prediction clients with a delay less than that of the prediction client, that is, during the pre-play process of the non-prediction clients, the second historical viewing trajectories of the viewing clients are obtained, which provides a reference for the perspective prediction of the prediction client, improves the accuracy of the predicted perspective of the prediction client, and further improves the accuracy of video encoding and video streaming. Description of the Drawings
[0020] Combined with the drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale.
[0021] Figure 1 is a schematic diagram of a live broadcast scenario provided by an embodiment of the present disclosure;
[0022] Figure 2 is a schematic flowchart of a method for streaming a live panoramic video provided by an embodiment of the present disclosure;
[0023] Figure 3 is a flowchart of a method for streaming a live panoramic video provided by an embodiment of the present disclosure;
[0024] Figure 4 is a flowchart of the streaming process provided by an embodiment of the present disclosure;
[0025] Figure 5 is a flowchart of perspective prediction provided by an embodiment of the present disclosure;
[0026] Figure 6 is a schematic structural diagram of a live panoramic video streaming device provided by an embodiment of the present disclosure;
[0027] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0028] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0029] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0030] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0031] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0032] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly stated in the context, it should be understood as "one or more".
[0033] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0034] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user in an appropriate manner and the user's authorization should be obtained in accordance with relevant laws and regulations.
[0035] For example, when responding to an active request from a user, a prompt message is sent to the user to clearly prompt that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that performs the operations of the technical solutions of the present disclosure based on the prompt message.
[0036] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0037] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manners of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0038] It can be understood that the data involved in the technical solutions (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.
[0039] See Figure 1 , Figure 1 is a schematic diagram of a live broadcast scenario provided by an embodiment of the present disclosure. The live broadcast scenario can be a VR live video scenario. Among them, the server receives the index video source uploaded by the anchor side, encodes the video source into live video streams with different resolutions, and pushes the live video streams to different clients. In the VR live video scenario, the video pushed to the client is a panoramic video. In order to simplify video encoding and accelerate video pushing, the viewing angle of the client is predicted, and targeted video encoding is performed according to the predicted viewing angle. Currently, the way of performing viewing angle prediction can be to perform prediction processing through the historical viewing trajectory of the predicted client, which has the problem of low viewing angle prediction accuracy. At the same time, different clients and different viewing users perform viewing angle prediction based on the same viewing angle prediction method, and it is impossible to perform differential viewing angle prediction for different clients and / or different viewing users.
[0040] During the video playback process, the parameters affecting the video playback quality include video picture quality and transmission delay. Different client users have different requirements for the playback of live videos. Exemplarily, the first type of users require low transmission delay for the video and are not sensitive to the picture quality. The second type of users require high video picture quality and are not sensitive to the transmission delay. The third type of users are not sensitive to both the transmission delay and the picture quality. For the above-mentioned first type of users and the second type of users, the delay of video streaming is different, that is, for the same live video, the delay of streaming to the first type of users is less than the delay of streaming to the second type of users.
[0041] According to the different viewing requirements of the viewing users, different streaming strategies and view prediction methods can be set to meet the viewing needs of different viewing users. In this embodiment, tags can be set for the client / viewing user. The tags include but are not limited to picture quality sensitive tags, delay sensitive tags, and other tags. Among them, in the case of setting a picture quality sensitive tag, it is required that the streamed video has high picture quality and is not sensitive to the transmission delay. In the case of setting a delay sensitive tag, it is required that the streamed video has low transmission delay and is not sensitive to the picture quality. That is, the delay of the streamed video of the client with a delay sensitive tag set is less than the delay of the streamed video of the client with a picture quality sensitive tag set. Correspondingly, in this embodiment, based on the historical viewing trajectory of the client with a delay sensitive tag set, the viewing angle of the client with a delay sensitive tag set can be predicted to improve the accuracy of view prediction of the client with a delay sensitive tag set.
[0042] In some embodiments, the tags set on the client can be set by the viewing user's operation. Exemplarily, a tag selection control can be set on the display interface of the client. The tag selection control can be a switching control, a drop-down selection control, etc., which is not limited herein. The tag selection control can select a picture quality sensitive tag, a delay sensitive tag, and other tags. Correspondingly, during the process of streaming to this client, the server can obtain the tag of this client and perform video streaming according to different streaming strategies based on the tag.
[0043] In some embodiments, tags can be set for the client based on the playback parameters of the videos historically played by the client, or tags can be set for the viewing user based on the playback parameters of the videos historically played by the client in a certain user login state. Optionally, the method for determining the tag of any client includes: obtaining the picture quality data and / or delay data of the videos historically played by the client within a preset time window; classifying the client based on the picture quality data and / or delay data of the videos historically played within the preset time window to obtain a classification result; setting a tag for the client based on the classification result, and the tag includes a picture quality sensitive tag and a delay sensitive tag.
[0044] Optionally, the tag of the client can be fixed and determined based on the historical played videos of the client. Specifically, it can be to obtain the historical played videos of the client within a preset time window and determine the corresponding tag. Optionally, the tag of the client can change according to the change of the logged-in user and is determined based on the historical played videos of the logged-in user on each client. Specifically, it can be to obtain the historical played videos of the logged-in user on the logged-in clients within a preset time window and determine the corresponding tag. The clients here include, but are not limited to, electronic devices with video playback functions such as mobile phones, tablets, PCs, VR devices, etc.
[0045] The relevant parameters of the historical played videos are stored in the client, and the server regularly obtains the relevant parameters of the historical played videos within a preset time window from the client to update the tags of the client or the logged-in user on the client. The historical played videos within the preset time window can be a preset number of historical played videos or the historical played videos within a preset time period before the moment of determining the tag.
[0046] The relevant parameters of the historical played videos include at least one of picture quality data and latency data, and the tag of the client or the logged-in user is determined based on at least one of the picture quality data and the latency data. Specifically, it can be to classify the client or the logged-in user based on at least one of the picture quality data and the latency data of the historical played videos to determine the classification result of the client or the logged-in user. Among them, the classification result includes at least latency-sensitive type and picture-quality-sensitive type. Correspondingly, set the latency-sensitive tag for the latency-sensitive client or logged-in user, and set the picture-quality-sensitive tag for the picture-quality-sensitive client or logged-in user. It can be understood that the classification result can also include other types except the latency-sensitive type and the picture-quality-sensitive type, and set other type tags accordingly. This other type can be the type that is not sensitive to both latency and picture quality.
[0047] Taking the picture quality data as an example, based on the picture quality data of the historical played videos within the preset time window, determine the video playback duration corresponding to each picture quality range; based on the video playback duration corresponding to each picture quality range, determine the first playback duration ratio corresponding to each picture quality range; determine the classification result based on the first playback duration ratio corresponding to each picture quality range. Among them, multiple picture quality ranges are preset, and the picture quality range at least includes a high picture quality range. For example, the high picture quality range can be that the picture quality data vqscore is greater than the picture quality threshold (such as 85). Determine the first playback duration ratio corresponding to each picture quality range. In particular, when the first playback duration ratio corresponding to the high picture quality range is greater than the first threshold, determine that the client or the logged-in user belongs to the picture-quality-sensitive type, and correspondingly, set the picture-quality-sensitive tag for the client or the logged-in user.
[0048] Taking the latency data as an example, the video playback duration corresponding to each latency range can be determined based on the latency data of the historical played videos within the preset time window; based on the video playback duration corresponding to each latency range, the second playback duration ratio corresponding to each latency range can be determined, and the classification result can be determined based on the second playback duration ratio corresponding to the latency range. Among them, multiple latency ranges are preset, and the latency range at least includes a low-latency range. For example, the low-latency range can be a latency less than the latency threshold (e.g., 100 ms). Determining the second playback duration ratio corresponding to each latency range, especially the second playback duration ratio corresponding to the low-latency range. In the case where the second playback duration ratio corresponding to the low-latency range is greater than the second threshold, it is determined that the client or the logged-in user belongs to the latency-sensitive category. Correspondingly, a latency-sensitive label is set for the client or the logged-in user.
[0049] In some embodiments, the client or the logged-in user is classified and labeled through the video quality data and latency data of the historical played videos. Specifically, based on the video quality data of the historical played videos, the video playback duration corresponding to each video quality range is determined; based on the video playback duration corresponding to each video quality range, the first playback duration ratio corresponding to each video quality range is determined; based on the latency data of the historical played videos, the video playback duration corresponding to each latency range is determined; based on the video playback duration corresponding to each latency range, the second playback duration ratio corresponding to each latency range is determined; based on the first playback duration ratio corresponding to each video quality range and the second playback duration ratio corresponding to each latency range, the client is classified to obtain a classification result.
[0050] Among them, different types correspond to different combinations of latency ranges and video quality ranges. Exemplarily, latency-sensitive types correspond to low-latency ranges and low-video-quality ranges, and video-quality-sensitive types correspond to high-video-quality ranges and high-latency ranges. For example, the low-latency range corresponding to the latency-sensitive type is that the latency is less than the first latency threshold (e.g., 100 ms), and the low-video-quality range can be that the video quality data is between the first video quality threshold (e.g., 70) and the second video quality threshold (e.g., 85). When the second playback duration ratio corresponding to the low-latency range and the first playback duration ratio corresponding to the low-video-quality range are both greater than the third threshold, it is determined that the client or the logged-in user belongs to the latency-sensitive type; or, when the playback duration ratio that simultaneously satisfies the low-latency range and the low-video-quality range is greater than the third threshold, it is determined that the client or the logged-in user belongs to the latency-sensitive type, and a latency-sensitive label is set. The high-video-quality range corresponding to the video-quality-sensitive type can be greater than the second video quality threshold, and the high-latency range can be that the latency data is greater than the second latency threshold (e.g., 500 ms). When the first playback duration ratio corresponding to the high-video-quality range and the second playback duration ratio corresponding to the high-latency range are both greater than the fourth threshold, it is determined that the client or the logged-in user belongs to the video-quality-sensitive type; or, when the playback duration ratio that simultaneously satisfies the high-video-quality range and the high-latency range is greater than the fourth threshold, it is determined that the client or the logged-in user belongs to the video-quality-sensitive type, and a video-quality-sensitive label is set.
[0051] In some embodiments, the first threshold, the second threshold, the third threshold, and the fourth threshold can be the same. For example, it can be 50%.
[0052] During the playback of the live panoramic video, for the client or the logged-in user of the latency-sensitive type, the historical viewing trajectory corresponding to the client or the logged-in user is used for perspective prediction to obtain a predicted perspective. After video encoding based on the predicted perspective, video streaming is performed to the corresponding client. The perspective prediction method is not elaborated here.
[0053] For the client or the logged-in user of the video-quality-sensitive type, during the playback of the live panoramic video, the latency is relatively long, especially compared with the client or the logged-in user of the latency-sensitive type. The historical viewing trajectory of the client of the latency-sensitive type and the historical viewing trajectory of the current client of the video-quality-sensitive type can be obtained for perspective prediction to obtain a predicted perspective. After video encoding based on the predicted perspective, video streaming is performed to the corresponding client.
[0054] See Figure 2 , Figure 2The figure is a schematic flowchart of a method for pushing live panoramic video provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of accurately pushing a live video for a client set with a picture quality sensitive label. This method can be executed by a pushing device for live panoramic video, and the device can be implemented in the form of software and / or hardware. Optionally, it is implemented through an electronic device, and the electronic device can be a mobile terminal, a PC, a server, etc. As Figure 2 shown, the method includes:
[0055] S110. When it is predicted that the client is set with a picture quality sensitive label, obtain the first historical viewing trajectory of the predicted client, and obtain the second historical viewing trajectories of multiple non-predicted clients set with a delay sensitive label, where the playback delay of the non-predicted clients is less than the playback delay of the predicted client.
[0056] S120. Based on the first historical viewing trajectory and the multiple second historical viewing trajectories, obtain the predicted viewing perspective of the predicted client.
[0057] S130. Perform panoramic video encoding based on the viewing perspective, and push the encoded video data to the predicted client.
[0058] Among them, the predicted client can be a client that requests a panoramic video from the server. The predicted client can be set with a label, and the label can be set for the predicted client, or for the user logged in to the client, or determined by collecting the setting operations of the current operating user, and no limitation is made thereto. In some embodiments, the predicted client sends a video request to the server, and the video request may carry a label. When the label corresponding to the predicted client is a picture quality sensitive label, the pushing method provided by the embodiment of the present disclosure is adopted to push the panoramic video to the predicted client. The non-predicted client can be a client sensitive to delay, and the viewing perspective is determined only through the historical viewing trajectory of the non-predicted client itself, and video encoding and pushing are performed.
[0059] The server communicates with each client. During the process of video streaming to each client, it receives the historical viewing trajectories feedback by each client. Specifically, it includes the first historical viewing trajectory of the predicted client and the second historical viewing trajectories of multiple non-predicted clients with latency-sensitive tags set. It can be understood that the above first historical viewing trajectory and second historical viewing trajectories are the viewing trajectories of the same panoramic video. In this embodiment, the panoramic video is a live panoramic video, and the playback latency of the non-predicted client is less than that of the predicted client. Correspondingly, the second historical viewing trajectory of the non-predicted client includes the historical viewing trajectory of some video segments where no video streaming is performed to the predicted client, and this part of the video segments is the video segments corresponding to the latency difference between the non-predicted client and the predicted client.
[0060] By obtaining the second historical viewing trajectories of multiple non-predicted clients, which include the historical associated trajectories corresponding to the latency difference between the predicted client and the non-predicted clients, based on the second historical viewing trajectories of the non-predicted clients and the first historical viewing trajectory of the predicted client, jointly perform perspective prediction on the predicted client, improve the accuracy of the predicted perspective, and further improve the accuracy of video streaming to the predicted client.
[0061] In this embodiment, perspective prediction can be performed based on a pre-trained perspective prediction model. For example, the first historical viewing trajectory and multiple said second historical viewing trajectories are input into the above perspective prediction model to obtain the predicted perspective output by the perspective prediction model. Among them, the perspective prediction model can be a neural network model, such as an LSTM (Long Short-Term Memory) model, etc. Here, the structural type and specific structure of the perspective prediction model are not limited.
[0062] In some embodiments, based on the first historical viewing trajectory and multiple said second historical viewing trajectories, obtaining the predicted viewing perspective of the predicted client includes: performing trajectory fusion processing on multiple said second historical viewing trajectories to obtain a fused viewing trajectory; based on the first historical viewing trajectory and the fused viewing trajectory, obtaining the predicted viewing perspective of the predicted client. Here, there are multiple second historical viewing trajectories. By fusing multiple second historical viewing trajectories to obtain a fused viewing trajectory, correspondingly, inputting the fused viewing trajectory and the first historical viewing trajectory into the perspective prediction model simplifies the input data of the perspective prediction model. It can be understood that the first historical viewing trajectory and the second historical viewing trajectories are the movement trajectories of the head during the process of the viewing user watching the live panoramic video, and this head movement trajectory is obtained by combining multiple head coordinate data.
[0063] In this embodiment, the fusion process of multiple second historical viewing trajectories may be a weighted process of multiple second historical viewing trajectories. Optionally, the trajectory fusion process of multiple second historical viewing trajectories to obtain a fused viewing trajectory includes: determining the weights of the second historical viewing trajectories, and based on the weights of the second historical viewing trajectories, performing a weighted fusion process on the multiple second historical viewing trajectories to obtain the fused viewing trajectory. In some embodiments, the weights of multiple second historical viewing trajectories may be the same or different. For each second historical viewing trajectory, the weight may be set according to the contribution degree of the second historical viewing trajectory to the viewing angle prediction.
[0064] Exemplarily, determining the weights of the second historical viewing trajectories includes: classifying the similarity of multiple second historical viewing trajectories to obtain at least two sets of trajectory collections; setting weights for the second historical viewing trajectories in each of the trajectory collections based on the number of similar trajectories in each of the trajectory collections. Optionally, clustering processing may be performed on the second historical viewing trajectories to obtain at least two classes, that is, at least two sets of trajectory collections, and the similarity of the second historical viewing trajectories in the same trajectory combination is higher than a preset threshold. Optionally, the similarity calculation may be performed on every two second historical viewing trajectories to obtain the similarity of every two second historical viewing trajectories. When the similarity is greater than the similarity threshold, it is determined that the above two second historical viewing trajectories belong to the same set of trajectory collections. The greater the number of similar trajectories in the same set of trajectory collections, the greater the contribution degree to the viewing angle prediction. Here, based on the number of similar trajectories in the trajectory collection, weights are set for the second historical viewing trajectories in each trajectory collection, and the weight of the second historical viewing trajectory is positively correlated with the number of similar trajectories in the trajectory collection to which it belongs.
[0065] Through the weights of the second historical viewing trajectories, weighted processing is performed on the second historical viewing trajectories to obtain a fused viewing trajectory. Through the fused viewing trajectory and the first historical viewing trajectory, the viewing angle prediction is performed on the prediction client. Through the predicted viewing angle, panoramic video encoding is performed to obtain video data suitable for the prediction client, and the video data is pushed to the prediction client so that the prediction client can display the panoramic video.
[0066] In some embodiments, video encoding may be performed based on the Equi-angular Cubemap (EAC) and offset cubic algorithms. Here, the video encoding method is not limited.
[0067] In the technical solution of the embodiments of the present disclosure, for a prediction client set with a picture quality sensitive tag, during the process of live panoramic video streaming, the first historical viewing trajectory of the prediction client and the second historical viewing trajectories of multiple non-prediction clients set with multiple delay sensitive tags are obtained. By obtaining the second historical viewing trajectories of multiple non-prediction clients with a delay less than that of the prediction client, that is, obtaining the second historical viewing trajectories of the viewing clients during the pre-playback process of the non-prediction clients, it provides a reference for the perspective prediction of the prediction client, improves the accuracy of the predicted perspective of the prediction client, and further improves the accuracy of video encoding and video streaming.
[0068] On the basis of the above embodiments, before obtaining the second historical viewing trajectories of multiple non-prediction clients set with delay sensitive tags, the method further includes: predicting the delay length of the prediction client based on the picture quality sensitive tag and network status data of the prediction client. Among them, the greater the delay length, the greater the delay difference from the non-prediction client. Correspondingly, the more historical viewing trajectories corresponding to the delay difference are obtained, providing more reference data for the perspective prediction of the prediction client. Correspondingly, the higher the accuracy of the predicted perspective. At the same time, the greater the delay length, the greater the transmission delay of the live panoramic video playback process of the prediction client.
[0069] By predicting the delay length of the prediction client, obtaining the second historical viewing trajectories of multiple non-prediction clients set with delay sensitive tags based on the delay length, and determining the timing of perspective prediction and video streaming for the prediction client based on the delay length. During the streaming process, the second historical viewing trajectories of the above delay length are cumulatively obtained, the initial perspective of the prediction client is predicted, video encoding is performed based on the initial perspective, and video streaming is performed for the prediction client. During the streaming process, the second historical viewing trajectories and the first historical viewing trajectories are continuously obtained in real time, and perspective prediction and video encoding are continued, and video streaming is continuously performed for the prediction client.
[0070] In this embodiment, by predicting the delay length based on the picture quality sensitive tag and network status data of the prediction client, where the network status data may be network speed data. For example, the network status data of the prediction client may be real-time data, or the stable network status data of the prediction client within a preset time range, such as average network status data. The worse the network status data, the greater the corresponding delay length.
[0071] Optionally, the quality-sensitive label of the prediction client can be a fixed label and can also be further divided into multiple sub-labels. The sub-label can be used to characterize the quality requirements within a preset time window. For example, the sub-label can be determined according to the quality data within the preset time window, such as any one of the mean, median, or mode of the quality data. The higher the quality requirements characterized by the sub-label, correspondingly, the longer the delay length.
[0072] Optionally, based on the quality-sensitive label and network status data of the prediction client, predicting the delay length of the prediction client includes: performing a prediction process on the quality-sensitive label and network status data of the prediction client based on a first prediction model to obtain the delay length of the prediction client. The first prediction model can be a machine learning model, such as a neural network model, etc. Correspondingly, inputting the quality-sensitive label and network status data of the above prediction client into the first prediction model to obtain the delay length output by the first prediction model.
[0073] Based on the above embodiments, before performing panoramic video encoding based on the viewing perspective, it further includes: predicting the encoding parameters of the prediction client based on the quality-sensitive label and network status data of the prediction client. In this embodiment, the offset cubic algorithm is used for video encoding, and the encoding parameters are the relevant parameters in the offset cubic algorithm. During the video encoding based on the offset cubic algorithm, upsampling encoding is performed on the main viewing direction (i.e., the predicted viewing direction), and the encoding parameters are data between -1 and 0. The smaller the encoding parameters, the smaller the range of the corresponding main viewing angle. Under the condition of ensuring the resolution and quality of the main viewing angle, the overall video bit rate is lower. By predicting the encoding parameters, panoramic video encoding can be performed based on the encoding parameters and the viewing perspective.
[0074] Optionally, predicting the encoding parameters of the prediction client based on the quality-sensitive label and network status data of the prediction client includes: performing a prediction process on the quality-sensitive label and network status data of the prediction client based on a second prediction model to obtain the encoding parameters of the prediction client. The second prediction model can be a machine learning model, such as a neural network model, etc.
[0075] In some embodiments, the first prediction model and the second prediction model can be the same prediction model, that is, inputting the quality-sensitive label and network status data of the prediction client into the above prediction model to obtain the encoding parameters and delay length output by the prediction model.
[0076] Based on the above embodiments, the method further includes: obtaining the actual viewing angle feedback by the prediction client, and updating and training the first prediction module and / or the second prediction model based on the actual viewing angle and the predicted viewing angle. Correspondingly, the encoding parameters and the delay length at the next moment are predicted based on the updated first prediction module and / or the second prediction model. By optimizing the first prediction module and / or the second prediction model in real time or at regular intervals, the prediction accuracy of the model is improved, the prediction accuracy of the encoding parameters and the delay length is improved, and correspondingly, the accuracy of video streaming is improved.
[0077] Based on the above embodiments, the present disclosure also provides a preferred example of a method for pushing live panoramic video. Refer to Figure 3 , Figure 3 which is a flowchart of a method for pushing live panoramic video provided by an embodiment of the present disclosure. Classify the client or the logged-in user on the client and set tags. According to the sensitivity to video quality, delay, etc. during video viewing, the client or the logged-in user on the client can be divided into three categories: User group A (i.e., delay-sensitive type): prefers low-delay transmission and is less sensitive to video quality;. User group B (i.e., video-quality-sensitive type): prefers high video quality and is less sensitive to transmission delay; User group C (i.e., other type): other users outside the above two groups A and B. Each type of client or logged-in user is respectively set with a tag, and according to the request sent by the client, the type to which the client belongs can be determined. During the historical video playback process of each client, a log is reported, and the log includes but is not limited to video quality data and delay data. The server determines the tags of the client or the logged-in user on the client according to the reported log and stores them.
[0078] Exemplarily, the video viewing data of the client or the logged-in user for a period of time (one week) is statistically analyzed, and the video viewing duration ratio under different combinations of delay and vqscore (video quality data, with a value range of 1-100, and the higher the score, the higher the video quality) is calculated. For example, if the viewing duration ratio of user X when vqscore>85 & delay>500ms exceeds 50%, it is considered that user X belongs to user group B; if the viewing duration ratio of user Y when 85>vqscore>70 & delay<100ms exceeds 50%, it is considered that user X belongs to user group A.
[0079] Meanwhile, during the playback of historical videos by each client, the uploaded logs also include network status data and historical viewing trajectories. Based on the tags of each client, the historical viewing trajectories of each client, and the network status data, the push stream + offset parameter configuration module is called for video push streaming. In the solution provided by the present disclosure, for users who are sensitive to latency and not sensitive to video quality, the video needs to be transmitted as soon as possible. Due to the need for quick response, it is impossible to collect sufficient perspective information of other users. Therefore, the accuracy of perspective prediction is relatively low, and thus the video quality cannot be improved by adjusting the offset coefficient. For users who are sensitive to video quality and not sensitive to latency, high-quality videos need to be transmitted to optimize the video playback quality. By collecting the viewing records of other users (non-prediction clients) using the time difference (i.e., latency difference), the result of perspective prediction can be improved. Meanwhile, by coordinating the adjustment of the encoding coefficient, it is possible to improve the video quality in the main perspective direction under the same bit rate.
[0080] See Figure 4 , Figure 4 is the flowchart of the push streaming process provided by the embodiments of the present disclosure. By obtaining the tags of the client or the logged-in user, and the network test result (i.e., network status data) of the prediction client, the above tags and network test results are input into the reinforcement learning model (i.e., the first prediction model or the second prediction model) to obtain the encoding offset coefficient (i.e., encoding parameter) and the number of seconds of latency acceptable to the user (i.e., latency length). Based on the latency length, perspective prediction is performed to obtain the main perspective direction (i.e., prediction perspective). Based on the prediction perspective and the encoding offset coefficient, panoramic video encoding is performed to obtain video data, which is then pushed to the prediction client.
[0081] See Figure 5 , Figure 5 is the flowchart of the perspective prediction provided by the embodiments of the present disclosure. Among them, the second historical viewing trajectories of multiple non-prediction clients are obtained, such as the historical viewing trajectories of non-prediction target users Y, Z, W, etc., and the first historical viewing trajectory of the prediction client, that is, the historical viewing trajectory of the prediction target user X. The second historical viewing trajectories of the non-prediction clients are weighted and averaged to obtain a fused viewing trajectory. Based on the fused viewing trajectory and the first historical viewing trajectory of the prediction client, the prediction perspective is jointly predicted. Specifically, the fused viewing trajectory and the first historical viewing trajectory of the prediction client can be input into a perspective prediction model (such as an LSTM model) to obtain the prediction perspective.
[0082] Figure 6 is a schematic structural diagram of a push streaming device for live panoramic video provided by the embodiments of the present disclosure, as Figure 6 shown, the device includes: a data acquisition module 210, a perspective prediction module 220, a video encoding module 230, and a video push streaming module 240.
[0083] A data acquisition module 210, configured to obtain a first historical viewing track of the prediction client when a picture quality sensitive tag is set on the prediction client, and obtain second historical viewing tracks of multiple non-prediction clients with a latency sensitive tag set, wherein the playback latency of the non-prediction clients is less than the playback latency of the prediction client;
[0084] A viewing angle prediction module 220, configured to obtain a predicted viewing angle of the prediction client based on the first historical viewing track and multiple second historical viewing tracks;
[0085] A video encoding module 230, configured to perform panoramic video encoding based on the viewing angle;
[0086] A video streaming module 240, configured to stream the encoded video data to the prediction client.
[0087] In the technical solution provided by the embodiments of the present disclosure, for a prediction client with a picture quality sensitive tag set, during the process of live panoramic video streaming, a first historical viewing track of the prediction client and second historical viewing tracks of multiple non-prediction clients with a latency sensitive tag set are obtained. By obtaining the second historical viewing tracks of multiple non-prediction clients with a latency less than that of the prediction client, that is, obtaining the second historical viewing tracks of the viewing clients during the pre-playback process of the non-prediction clients, it provides a reference for the viewing angle prediction of the prediction client, improves the accuracy of the predicted viewing angle of the prediction client, and further improves the accuracy of video encoding and video streaming.
[0088] Based on the above embodiments, optionally, the apparatus further includes a tag determination module, including:
[0089] A video parameter acquisition unit, configured to acquire picture quality data and / or latency data of the historical video played by the client within a preset time window;
[0090] A classification unit, configured to classify the client based on the picture quality data and / or latency data of the historical video played within the preset time window to obtain a classification result;
[0091] A tag setting unit, configured to set tags for the client based on the classification result, where the tags include a picture quality sensitive tag and a latency sensitive tag.
[0092] Optionally, the classification unit is configured to:
[0093] Based on the picture quality data of the historical video played within the preset time window, determine the video playback duration corresponding to each picture quality range; based on the video playback duration corresponding to each picture quality range, determine the first playback duration ratio corresponding to each picture quality range;
[0094] And / or, based on the latency data of historical played videos within the preset time window, determine the video playing duration corresponding to each latency range; based on the video playing duration corresponding to each latency range, determine the second playing duration ratio corresponding to each latency range.
[0095] Classify the client based on the first playing duration ratio corresponding to each picture quality range and / or the second playing duration ratio corresponding to each latency range, to obtain a classification result.
[0096] Based on the above embodiments, optionally, the perspective prediction module 220 includes:
[0097] A trajectory fusion unit, configured to perform trajectory fusion processing on multiple second historical viewing trajectories to obtain a fused viewing trajectory.
[0098] A perspective prediction unit, configured to obtain the predicted viewing perspective of the predicted client based on the first historical viewing trajectory and the fused viewing trajectory.
[0099] Optionally, the trajectory fusion unit is configured to: determine the weight of each second historical viewing trajectory, and perform weighted fusion processing on multiple second historical viewing trajectories based on the weight of the second historical viewing trajectory to obtain the fused viewing trajectory.
[0100] Optionally, the trajectory fusion unit is configured to: perform trajectory similarity classification on multiple second historical viewing trajectories to obtain at least two groups of trajectory sets; set weights for the second historical viewing trajectories in each trajectory set based on the number of similar trajectories in each trajectory set.
[0101] Based on the above embodiments, optionally, the apparatus further includes: a latency length determination module, configured to predict the latency length of the predicted client based on the picture quality sensitive label and network status data of the predicted client before obtaining the second historical viewing trajectories of multiple non-predicted clients set with latency sensitive labels.
[0102] Correspondingly, the data acquisition module 210 is configured to obtain the second historical viewing trajectories of multiple non-predicted clients set with latency sensitive labels based on the latency length.
[0103] Based on the above embodiments, optionally, the apparatus further includes: an encoding parameter prediction module, configured to: before performing panoramic video encoding based on the viewing perspective, predict the encoding parameters of the predicted client based on the picture quality sensitive label and network status data of the predicted client.
[0104] The video encoding module 230 is configured to: perform panoramic video encoding based on the encoding parameters and the viewing perspective.
[0105] Optionally, the latency length determination module is configured to: based on a first prediction model, perform prediction processing on the picture quality sensitive label and network state data of the prediction client to obtain the latency length of the prediction client.
[0106] The encoding parameter prediction module is configured to: based on a second prediction model, perform prediction processing on the picture quality sensitive label and network state data of the prediction client to obtain the encoding parameters of the prediction client.
[0107] Based on the above embodiments, optionally, the apparatus further includes: a model optimization module, configured to obtain the actual viewing perspective fed back by the prediction client, and based on the actual viewing perspective and the predicted viewing perspective, update and train the first prediction module and / or the second prediction model.
[0108] The live panoramic video streaming apparatus provided by the embodiments of the present disclosure can execute the live panoramic video streaming method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.
[0109] It should be noted that the various units and modules included in the above apparatus are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.
[0110] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Referring below to Figure 7 , which shows a schematic structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure (such as Figure 7 the terminal device or server in). The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0111] Such as Figure 7As shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An editing / output (I / O) interface 505 is also connected to the bus 504.
[0112] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 an electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0113] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0114] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0115] The electronic device provided by the embodiment of the present disclosure and the method for pushing a live panoramic video provided by the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment may be referred to in the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0116] An embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the method for pushing a live panoramic video provided in the above embodiment is implemented.
[0117] It should be noted that the computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0118] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (for example, a communication network). Examples of the communication network include a local area network ("LAN"), a wide area network ("WAN"), the Internet (for example, the Internet), and a peer-to-peer network (for example, an ad hoc peer-to-peer network), as well as any currently known or future-developed network.
[0119] The above computer-readable medium may be included in the above electronic device; or it may exist separately and not be assembled into the electronic device.
[0120] The above computer-readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to:
[0121] The above computer-readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to: when a picture quality sensitive tag is set in the prediction client, obtain a first historical viewing track of the prediction client, and obtain second historical viewing tracks of a plurality of non-prediction clients with a latency sensitive tag set, wherein a playback latency of the non-prediction client is less than a playback latency of the prediction client; based on the first historical viewing track and the plurality of second historical viewing tracks, obtain a predicted viewing perspective of the prediction client; perform panoramic video encoding based on the viewing perspective, and push the encoded video data to the prediction client.
[0122] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0124] The units involved in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring at least two Internet protocol addresses".
[0125] The functions described above in this article can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0126] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disc read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0127]
In the detailed implementation section, after the full text ends, please repeat all the content to be protected in the form of claims as follows:
[0128] According to one or more embodiments of the present disclosure,
Example 1
[0129] When a picture quality sensitive label is set on the prediction client, acquiring the first historical viewing track of the prediction client, and acquiring the second historical viewing tracks of multiple non-prediction clients with a latency sensitive label set, where the playback latency of the non-prediction clients is less than the playback latency of the prediction client; obtaining the predicted viewing perspective of the prediction client based on the first historical viewing track and the multiple second historical viewing tracks; performing panoramic video encoding based on the viewing perspective, and pushing the encoded video data to the prediction client.
[0130] According to one or more embodiments of the present disclosure, [Example 2] provides a method for pushing a live panoramic video in Example 1, further including:
[0131] A method for determining a tag of any client includes: obtaining the video quality data and / or latency data of the historical video played by the client within a preset time window; classifying the client based on the video quality data and / or latency data of the historical video played within the preset time window to obtain a classification result; setting a tag for the client based on the classification result, where the tag includes a video quality sensitive tag and a latency sensitive tag.
[0132] According to one or more embodiments of the present disclosure, [Example 3] provides a method for pushing a live panoramic video in Example 1, further including:
[0133] The classifying the client based on the video quality data and latency data of the historical video played within the preset time window to obtain a classification result includes: determining the video playback duration corresponding to each video quality range based on the video quality data of the historical video played within the preset time window; determining the first playback duration ratio corresponding to each video quality range based on the video playback duration corresponding to each video quality range;
[0134] And / or, determining the video playback duration corresponding to each latency range based on the latency data of the historical video played within the preset time window; determining the second playback duration ratio corresponding to each latency range based on the video playback duration corresponding to each latency range;
[0135] Classifying the client based on the first playback duration ratio corresponding to each video quality range and / or the second playback duration ratio corresponding to each latency range to obtain a classification result.
[0136] According to one or more embodiments of the present disclosure, [Example 4] provides a method for pushing a live panoramic video in Example 1, further including:
[0137] Obtaining the predicted viewing angle of the predicted client based on the first historical viewing trajectory and multiple second historical viewing trajectories includes: performing trajectory fusion processing on the multiple second historical viewing trajectories to obtain a fused viewing trajectory; obtaining the predicted viewing angle of the predicted client based on the first historical viewing trajectory and the fused viewing trajectory.
[0138] According to one or more embodiments of the present disclosure, [Example 5] provides a method for pushing a live panoramic video in Example 1, further including:
[0139] Performing trajectory fusion processing on the multiple second historical viewing trajectories to obtain a fused viewing trajectory includes: determining the weights of the second historical viewing trajectories, and performing weighted fusion processing on the multiple second historical viewing trajectories based on the weights of the second historical viewing trajectories to obtain the fused viewing trajectory.
[0140] According to one or more embodiments of the present disclosure, [Example Six] provides a method for pushing a live panoramic video in Example One, further including:
[0141] Determining the weights of the second historical viewing trajectories includes: performing trajectory similarity classification on the multiple second historical viewing trajectories to obtain at least two sets of trajectory collections; setting weights for the second historical viewing trajectories in each of the trajectory collections based on the number of similar trajectories in each of the trajectory collections.
[0142] According to one or more embodiments of the present disclosure, [Example Seven] provides a method for pushing a live panoramic video in Example One, further including:
[0143] Before obtaining the second historical viewing trajectories of multiple non-predictive clients with delay-sensitive tags set, the method further includes: predicting the delay length of the predictive client based on the picture quality sensitive tag and network status data of the predictive client;
[0144] Correspondingly, obtaining the second historical viewing trajectories of multiple non-predictive clients with delay-sensitive tags set includes: obtaining the second historical viewing trajectories of the multiple non-predictive clients with delay-sensitive tags set based on the delay length.
[0145] According to one or more embodiments of the present disclosure, [Example Eight] provides a method for pushing a live panoramic video in Example One, further including:
[0146] Before performing panoramic video encoding based on the viewing angle, it further includes: predicting the encoding parameters of the predictive client based on the picture quality sensitive tag and network status data of the predictive client;
[0147] Correspondingly, performing panoramic video encoding based on the viewing angle includes: performing panoramic video encoding based on the encoding parameters and the viewing angle.
[0148] According to one or more embodiments of the present disclosure, [Example Nine] provides a method for pushing a live panoramic video in Example One, further including:
[0149] Predicting the latency length of the prediction client based on the picture quality sensitive label and network status data of the prediction client includes: predicting and processing the picture quality sensitive label and network status data of the prediction client based on a first prediction model to obtain the latency length of the prediction client;
[0150] Or,
[0151] Predicting the encoding parameters of the prediction client based on the picture quality sensitive label and network status data of the prediction client includes: predicting and processing the picture quality sensitive label and network status data of the prediction client based on a second prediction model to obtain the encoding parameters of the prediction client.
[0152] According to one or more embodiments of the present disclosure, [Example Ten] provides a method for pushing a live panoramic video in Example One, further including:
[0153] The method further includes: obtaining the actual viewing angle feedback by the prediction client, and updating and training the first prediction module and / or the second prediction model based on the actual viewing angle and the predicted viewing angle.
[0154] According to one or more embodiments of the present disclosure, [Example Eleven] provides a device for pushing a live panoramic video in Example Eleven, further including:
[0155] A data acquisition module, configured to obtain a first historical viewing trajectory of the prediction client when a picture quality sensitive label is set on the prediction client, and obtain a second historical viewing trajectory of a plurality of non-prediction clients with a latency sensitive label set, where the playback latency of the non-prediction client is less than the playback latency of the prediction client;
[0156] A viewing angle prediction module, configured to obtain a predicted viewing angle of the prediction client based on the first historical viewing trajectory and a plurality of the second historical viewing trajectories;
[0157] A video encoding module, configured to perform panoramic video encoding based on the viewing angle;
[0158] A video pushing module, configured to push the encoded video data to the prediction client.
[0159] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0160] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0161] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for pushing a live panoramic video stream, characterized in that, Including: When a video quality sensitive tag is set on the prediction client, obtaining the first historical viewing trajectory of the prediction client, and obtaining the second historical viewing trajectories of multiple non-prediction clients with a latency sensitive tag set, where the playback latency of the non-prediction client is less than the playback latency of the prediction client; Based on the first historical viewing trajectory and the multiple second historical viewing trajectories, obtaining the predicted viewing perspective of the prediction client; Performing panoramic video encoding based on the viewing perspective and pushing the encoded video data to the prediction client; Among them, the method for determining the tag of any client includes: Obtaining the video quality data and / or latency data of the historical videos played by the client within a preset time window; Classifying the client based on the video quality data and / or latency data of the historical videos played within the preset time window to obtain a classification result; Setting a tag for the client based on the classification result, where the tag includes a video quality sensitive tag and a latency sensitive tag.
2. The method according to claim 1, characterized in that The classifying the client based on the video quality data and latency data of the historical videos played within the preset time window to obtain a classification result includes: Based on the video quality data of the historical videos played within the preset time window, determining the video playback duration corresponding to each video quality range; based on the video playback duration corresponding to each video quality range, determining the first playback duration ratio corresponding to each video quality range; And / or, based on the latency data of the historical videos played within the preset time window, determining the video playback duration corresponding to each latency range; based on the video playback duration corresponding to each latency range, determining the second playback duration ratio corresponding to each latency range; Classifying the client based on the first playback duration ratio corresponding to each video quality range and / or the second playback duration ratio corresponding to each latency range to obtain a classification result.
3. The method according to claim 1, wherein The obtaining the predicted viewing perspective of the prediction client based on the first historical viewing trajectory and the multiple second historical viewing trajectories includes: Performing trajectory fusion processing on the multiple second historical viewing trajectories to obtain a fused viewing trajectory; Based on the first historical viewing trajectory and the fused viewing trajectory, obtaining the predicted viewing perspective of the prediction client.
4. The method according to claim 3, characterized in that The performing trajectory fusion processing on the multiple second historical viewing trajectories to obtain a fused viewing trajectory includes: Determining the weight of each second historical viewing trajectory, and based on the weight of the second historical viewing trajectory, performing weighted fusion processing on the multiple second historical viewing trajectories to obtain the fused viewing trajectory.
5. The method according to claim 4, wherein The determining the weight of each second historical viewing trajectory includes: Performing trajectory similarity classification on the multiple second historical viewing trajectories to obtain at least two groups of trajectory sets; Setting the weight for the second historical viewing trajectory in each trajectory set based on the number of similar trajectories in each trajectory set.
6. The method according to claim 1, characterized in that, Before obtaining the second historical viewing trajectories of the multiple non-prediction clients with a latency sensitive tag set, the method further includes: Predict the latency length of the prediction client based on the picture quality sensitive label and network status data of the prediction client; Correspondingly, obtaining the second historical viewing trajectories of multiple non-prediction clients with latency sensitive labels includes: Obtain the second historical viewing trajectories of multiple non-prediction clients with latency sensitive labels based on the latency length.
7. The method according to claim 1, wherein Before panoramic video encoding based on the viewing perspective, it further includes: Predict the encoding parameters of the prediction client based on the picture quality sensitive label and network status data of the prediction client; Correspondingly, panoramic video encoding based on the viewing perspective includes: Perform panoramic video encoding based on the encoding parameters and the viewing perspective.
8. The method according to claim 6 or 7, characterized in that, The predicting the latency length of the prediction client based on the picture quality sensitive label and network status data of the prediction client includes: performing prediction processing on the picture quality sensitive label and network status data of the prediction client based on a first prediction model to obtain the latency length of the prediction client; Or, The predicting the encoding parameters of the prediction client based on the picture quality sensitive label and network status data of the prediction client includes: performing prediction processing on the picture quality sensitive label and network status data of the prediction client based on a second prediction model to obtain the encoding parameters of the prediction client.
9. The method according to claim 8, characterized in that, The method further includes: Obtain the actual viewing perspective feedback by the prediction client, and update and train the first prediction model and / or the second prediction model based on the actual viewing perspective and the predicted viewing perspective.
10. A live panoramic video streaming device, characterized in that, It includes: A data acquisition module, configured to obtain the first historical viewing trajectory of the prediction client when the prediction client is set with a picture quality sensitive label, and obtain the second historical viewing trajectories of multiple non-prediction clients with latency sensitive labels, where the playback latency of the non-prediction client is less than the playback latency of the prediction client; A viewing perspective prediction module, configured to obtain the predicted viewing perspective of the prediction client based on the first historical viewing trajectory and multiple second historical viewing trajectories; A video encoding module, configured to perform panoramic video encoding based on the viewing perspective; A video streaming module, configured to stream the encoded video data to the prediction client; The device further includes a label determination module, and the label determination module includes: A video parameter acquisition unit, configured to acquire the picture quality data and / or latency data of the historical videos played by the client within a preset time window; A classification unit, configured to classify the client based on the picture quality data and / or latency data of the historical videos played within the preset time window to obtain a classification result; A label setting unit, configured to set labels for the client based on the classification result, and the labels include picture quality sensitive labels and latency sensitive labels.
11. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for pushing a live panoramic video as described in any one of claims 1-9.
12. A storage medium containing computer-executable instructions that, when executed by a computer processor, are used to execute the method for pushing a live panoramic video as described in any one of claims 1-9.
Citation Information
Patent Citations
View angle prediction method and device, equipment and storage medium
CN114827750A
Vehicle-mounted live streaming method and apparatus
WO2022166263A1