A Collaborative Live Streaming Interactive Data Processing Method, System, Device and Medium
By collecting multi-dimensional time series and artificial intelligence prediction models to optimize live broadcast data flow, the problem of unbalanced video stream quality and out-of-synchronization of playback in collaborative live broadcasts of multiple anchors is solved, and the synchronization of multi-anchors' video timing and the reduction of interaction delay is achieved, improving the viewer's viewing experience.
Patent Information
- Application Number
- CN202510477355.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-16
AI Technical Summary
During the collaborative live broadcast of multiple anchors, there are problems such as unbalanced video streaming quality, high interaction delay and out-synchronized playback content, especially when network bandwidth fluctuates or frequent audience interaction, which affects the viewing experience of the audience.
By collecting data on the anchor side and platform side, a multi-dimensional time series is generated, and the artificial intelligence prediction model is used to predict bandwidth status, packet loss rate and interaction delay, the live broadcast data flow is dynamically optimized, and the video frame time stamp synchronization is performed to achieve accurate alignment of the video timing of multiple anchors.
It effectively reduces the interaction delay and lag in the collaborative live broadcast of multiple anchors, improves the viewing experience of the audience, and ensures the synchronous playback of audio and video content of multiple anchors.
Smart Images

Figure CN120017874B_ABST
Abstract
Description
Technical Field
[0001] This application relates to data processing technologies, and in particular to a collaborative live broadcast interaction data processing method, system, device, and medium. Background Art
[0002] With the development of multi-anchor collaborative live broadcasts, more and more live broadcast scenarios require multiple anchors to connect and interact simultaneously and jointly output content. However, due to factors such as different geographical distributions, inconsistent network environments, and complex viewer interaction behaviors, it often leads to problems such as some anchors experiencing lag, high latency, or out-of-sync playback content, seriously affecting the viewing experience of the overall collaborative live broadcast for viewers.
[0003] In the prior art during the process of multi-anchor collaborative live broadcasts, it is often unable to accurately coordinate the video stream quality of each anchor, and there is also a lack of effective interaction optimization and video time sequence synchronization mechanisms, resulting in obvious playback differences between different anchors. Especially in the case of bandwidth fluctuations or frequent viewer interactions, the problem is more prominent. Summary of the Invention
[0004] This application provides a collaborative live broadcast interaction data processing method, system, device, and medium to solve the problems of the prior art.
[0005] In a first aspect, this application provides a collaborative live broadcast interaction data processing method, including:
[0006] Collecting collaborative live broadcast data, collecting anchor-side data and platform-side data, where the anchor-side data includes bandwidth, interaction latency, packet loss rate, jitter, and video frame timestamps, and the platform-side data includes the number of anchors, interaction frequency, and the number of viewers, and generating a multi-dimensional time series based on the anchor-side data and the platform-side data;
[0007] Data prediction based on an artificial intelligence prediction model, where the artificial intelligence prediction model predicts the bandwidth state prediction sequence, packet loss rate prediction sequence, and interaction latency prediction sequence of each anchor within a first preset time based on the multi-dimensional time series;
[0008] Dynamic optimization of the live broadcast data flow, dynamically optimizing the live broadcast data flow of each anchor according to the live broadcast data flow adaptive strategy based on the bandwidth state prediction sequence, packet loss rate prediction sequence, and interaction latency prediction sequence;
[0009] Optimization of collaborative live broadcast interaction, constructing an interaction experience function based on the interaction latency prediction sequence, globally evaluating the interaction latency of multiple anchors, and dynamically adjusting the interaction between the anchor and the viewer according to the collaborative live broadcast interaction adjustment strategy;
[0010] Multi-anchor video timing synchronization, determining a synchronization reference time based on the video frame timestamps of all the anchors, and performing corresponding video synchronization processing according to the time differences between the video frame timestamps of each anchor and the synchronization reference time.
[0011] In a possible design, the live data stream adaptive strategy includes a dynamic bitrate adjustment strategy and a resolution and frame rate adjustment strategy;
[0012] The dynamic bitrate adjustment strategy is to calculate and adjust the dynamic bitrate of each anchor in real time according to the bandwidth fluctuation prediction sequence, packet loss rate prediction sequence, and interaction delay prediction sequence;
[0013] The resolution and frame rate adjustment strategy is to adjust the resolution and frame rate of each anchor in real time according to the dynamic bitrate, including:
[0014] When the dynamic bitrate is greater than or equal to 6 Mbps, the resolution is 4K and the frame rate is 60 FPS;
[0015] When the dynamic bitrate is greater than or equal to 4 Mbps and less than 6 Mbps, the resolution is 2K and the frame rate is 60 FPS;
[0016] When the dynamic bitrate is greater than or equal to 2.5 Mbps and less than 4 Mbps, the resolution is 1080P and the frame rate is 60 FPS;
[0017] When the dynamic bitrate is greater than or equal to 1.5 Mbps and less than 2.5 Mbps, the resolution is 1080P and the frame rate is 30 FPS;
[0018] When the dynamic bitrate is greater than or equal to 1.0 Mbps and less than 1.5 Mbps, the resolution is 720P and the frame rate is 30 FPS;
[0019] When the dynamic bitrate is greater than or equal to 0.6 Mbps and less than 1.0 Mbps, the resolution is 480P and the frame rate is 30 FPS;
[0020] When the dynamic bitrate is greater than or equal to 0.4 Mbps and less than 0.6 Mbps, the resolution is 360P and the frame rate is 30 FPS.
[0021] In a possible design, the interaction experience function is the trigger basis for the collaborative live interaction adjustment strategy;
[0022] When the interaction experience function is greater than the interaction optimization trigger threshold, execute the collaborative live interaction adjustment strategy;
[0023] The interaction optimization trigger threshold is used to determine whether the interaction experience function is within an acceptable range.
[0024] In a possible design, the collaborative live interaction adjustment strategy includes:
[0025] Calculate the target interaction delay threshold for each host;
[0026] Compare the interaction delay prediction sequence of each host with the corresponding target interaction delay threshold;
[0027] When the interaction delay prediction sequence of the host is greater than the corresponding target interaction delay threshold, determine that the host is an object for interaction optimization;
[0028] Perform interaction optimization on the host determined to be an object for interaction optimization. The interaction optimization includes reducing the barrage sending frequency, lowering the comment interaction frequency, simplifying the reward animation special effects, and combining the display of gift special effects.
[0029] In a possible design, the target interaction delay threshold is a dynamic threshold.
[0030] In a possible design, the multi-host video timing synchronization includes:
[0031] Generate a video frame timestamp sequence for each host based on the video frame timestamps collected within a second preset time;
[0032] According to the video frame timestamp sequences of all hosts, adopt the maximum value strategy to determine the synchronization reference time;
[0033] Calculate the time difference between the video frame timestamp of each host and the synchronization reference time;
[0034] According to the time difference, perform frame interpolation operations or frame dropping operations on the video frames of each host respectively, including:
[0035] When the time difference is greater than 0, perform a frame interpolation operation on the video frame of the host;
[0036] When the time difference is less than 0, perform a frame dropping operation on the video frame of the host;
[0037] When the time difference is 0, the video frame of the host is not processed.
[0038] In a possible design, the frame interpolation operation is a frame interpolation algorithm based on an artificial intelligence prediction model. By inputting adjacent video frames into the artificial intelligence prediction model, intermediate transition frames are predicted and generated.
[0039] In a second aspect, the present application provides a collaborative live interaction data processing system, including:
[0040] The collaborative live broadcast data acquisition module collects the data on the host side and the data on the platform side. The data on the host side includes bandwidth, interaction latency, packet loss rate, jitter, and video frame timestamps. The data on the platform side includes the number of hosts, interaction frequency, and the number of audiences. A multi-dimensional time series is generated based on the data on the host side and the data on the platform side;
[0041] The data prediction module based on the artificial intelligence prediction model. The artificial intelligence prediction model is based on the multi-dimensional time series and predicts the bandwidth status prediction sequence, packet loss rate prediction sequence, and interaction latency prediction sequence of each host within the first preset time;
[0042] The live broadcast data flow dynamic optimization module dynamically optimizes the live broadcast data flow of each host according to the live broadcast data flow adaptive strategy based on the bandwidth status prediction sequence, packet loss rate prediction sequence, and interaction latency prediction sequence;
[0043] The collaborative live broadcast interaction optimization module constructs an interaction experience function based on the interaction latency prediction sequence, globally evaluates the multi-host interaction latency, and dynamically adjusts the interaction between the host and the audience according to the collaborative live broadcast interaction adjustment strategy;
[0044] The multi-host video time series synchronization module determines the synchronization reference time based on the video frame timestamps of all hosts, and performs corresponding video synchronization processing respectively according to the time difference between the video frame timestamp of each host and the synchronization reference time.
[0045] In a third aspect, the present application provides an electronic device, including:
[0046] A processor; and,
[0047] A memory for storing the executable instructions of the processor;
[0048] Wherein, the processor is configured to execute any possible method described in the first aspect by executing the executable instructions.
[0049] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement any possible method described in the first aspect.
[0050] The collaborative live broadcast interaction data processing method, system, device and medium provided by the present application are based on the collected data on the host side and the data on the platform side, intelligently predict the live broadcast data flow through an artificial intelligence prediction model, and realize the dynamic optimization of the live broadcast data flow and the optimization of collaborative live broadcast interaction based on the prediction results, effectively reducing the interaction latency and jitter problems in the process of multi-host collaborative live broadcast. At the same time, through the multi-host video time series synchronization, the accurate alignment of the multi-host audio and video content is realized, further improving the viewing experience of the audience. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0052] Figure 1 It is a schematic flowchart of a collaborative live broadcast interaction data processing method shown according to an exemplary embodiment of the present application;
[0053] Figure 2 It is a schematic flowchart of a collaborative live broadcast interaction adjustment strategy shown according to an exemplary embodiment of the present application;
[0054] Figure 3 It is a schematic flowchart of multi-anchor video time sequence synchronization shown according to an exemplary embodiment of the present application;
[0055] Figure 4 It is a schematic structural diagram of a collaborative live broadcast interaction data processing system shown according to an exemplary embodiment of the present application;
[0056] Figure 5 It is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of the present application.
[0057] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0059] Figure 1 It is a schematic flowchart of a collaborative live broadcast interaction data processing method shown according to an exemplary embodiment of the present application. As Figure 1 shown, the method provided in this embodiment includes:
[0060] Step S101: Collect collaborative live broadcast data, collect data on the anchor side and the platform side. The data on the anchor side includes bandwidth, interaction latency, packet loss rate, jitter, and video frame timestamps. The data on the platform side includes the number of anchors, interaction frequency, and the number of audiences. Generate a multi-dimensional time series based on the data on the anchor side and the platform side.
[0061] In this step, based on the collected data of the host side and the platform side, corresponding multi-dimensional time series can be generated. The following is an example to illustrate the correspondence between the data of the host side and the platform side and the multi-dimensional time series.
[0062] For example, the collected data of the host side and the platform side are shown in the following table:
[0063] Table 1 Data of the host side and the platform side
[0064]
[0065] The generated multi-dimensional time series are shown as follows:
[0066]
[0067] Among them, X is the multi-dimensional time series, is the feature vector of the i-th host at time t, is:
[0068]
[0069] The input feature vector corresponding to Table 1 is:
[0070]
[0071] The multi-dimensional time series X corresponding to Table 1 is:
[0072]
[0073] Step S102: Data prediction based on the artificial intelligence prediction model. The artificial intelligence prediction model is based on the multi-dimensional time series and predicts the bandwidth state prediction sequence, packet loss rate prediction sequence, and interaction delay prediction sequence of each host within the first preset time.
[0074] In this step, the artificial intelligence prediction model is constructed based on the long short-term memory network (LSTM) and trained on the historically collected co-live dataset using the supervised learning method. The co-live dataset used is the multi-dimensional time series jointly composed of the data of the host side and the platform side. The data of the host side includes bandwidth, interaction delay, packet loss rate, and jitter; the data of the platform side includes the number of hosts, interaction frequency, and the number of audiences.
[0075] The data types and structures input during the model training phase and the prediction phase are consistent, and are all used to model the evolution law of each anchor's network state over time. The artificial intelligence prediction model learns the temporal dependence features of this multi-dimensional time series to predict the anchor bandwidth state, interaction delay, and packet loss rate within the first preset time, and outputs the corresponding prediction sequence results. The roles of the collected data in model training and prediction are as follows:
[0076] Bandwidth: Continuously collect bandwidth values within a fixed window time T to construct the historical bandwidth time series of each anchor. The historical bandwidth time series is one of the input feature dimensions of the artificial intelligence prediction model, which is used for the artificial intelligence prediction model to learn the dynamic evolution law of the bandwidth state over time, and then predict the bandwidth state of each anchor within the first preset time t, and output the bandwidth state prediction sequence of each anchor. The prediction results are used to dynamically adjust the bitrate, resolution, and frame rate on the anchor side in real time.
[0077] Interaction delay: Continuously collect interaction delay values within a fixed window time T. The interaction delay is the round-trip response delay from the anchor side to the platform side, and construct the historical interaction delay sequence of each anchor. The historical interaction delay sequence is one of the input feature dimensions of the artificial intelligence prediction model, which is used for the artificial intelligence prediction model to learn the characteristics of network transmission delay fluctuations, and then predict the interaction delay state of each anchor within the first preset time t, and output the interaction delay prediction sequence of each anchor. The prediction results are used to dynamically adjust the bitrate, resolution, and frame rate on the anchor side in real time, and as a collaborative live interaction adjustment strategy to collaboratively adjust the interaction between each anchor and the audience.
[0078] Packet loss rate: Continuously collect the packet loss rate within a fixed window time T to construct the historical packet loss rate time series of each anchor. The packet loss rate is used to characterize the loss ratio of data packets during the network transmission process on the anchor side, and is an important indicator for measuring network stability. The historical packet loss rate time series is one of the input feature dimensions of the artificial intelligence prediction model, which is used for the artificial intelligence prediction model to learn the evolution characteristics of network transmission anomalies and packet loss fluctuations, and then predict the packet loss state of each anchor within the first preset time t, and output the packet loss rate prediction sequence of each anchor. The prediction results are used to dynamically adjust the bitrate, resolution, and frame rate on the anchor side in real time.
[0079] Jitter: Collect continuous network jitter values within a fixed time window T to construct the historical jitter time series for each live streamer. Jitter refers to the degree of fluctuation in the arrival times between consecutive data packets and is an important indirect indicator of network stability. The historical jitter time series, as one of the input features of the artificial intelligence prediction model, plays an auxiliary role in modeling during the training and prediction processes of the artificial intelligence prediction model for generating the bandwidth status prediction series, packet loss rate prediction series, and interaction delay prediction series, helping to improve the overall prediction accuracy and stability. For example, when the network is in the initial stage of congestion and the bandwidth and packet loss rate have not changed significantly, jitter usually increases first. The model can use the historical jitter time series to capture potential network fluctuation trends, thereby adjusting the output prediction results in advance and improving the timeliness and accuracy of the prediction.
[0080] Number of live streamers: Collect the statistical value of the number of live streamers within a fixed time window T to construct the historical number of live streamers time series. The number of live streamers is obtained through real-time statistics on the platform side and is used to represent the total number of live streamers participating in the interaction in the current co-live broadcast scenario. The historical number of live streamers time series is not used as an input feature dimension during the training and prediction processes of the artificial intelligence prediction model but as a control parameter to guide the artificial intelligence prediction model to determine the number of live streamers for which the current prediction task needs to be performed, that is, to control the number of model instantiations.
[0081] Interaction frequency: Collect the number of interaction messages per unit time for each live streamer within a fixed time window T to construct the historical interaction frequency time series for each live streamer. The interaction frequency reflects the interaction activity between the live streamer and the audience and is a key factor affecting the load of the interaction channel. The historical interaction frequency time series, as one of the input feature dimensions of the artificial intelligence prediction model, plays an auxiliary role in modeling during the process of generating the interaction delay prediction series, improving the accuracy of the interaction delay prediction. When the interaction frequency increases, the transmission load on the live streamer side increases accordingly, which may lead to fluctuations in the response of the feedback path. The artificial intelligence prediction model can identify potential trends of a sharp increase in interaction delay based on this.
[0082] Number of viewers: Collect the access statistics of the number of viewers in each live streamer's live broadcast room within a fixed time window T to construct the historical number of viewers time series for each live streamer. The number of viewers, as one of the metrics for measuring the system transmission load, affects the stability of the network transmission link. The historical number of viewers time series, as one of the input feature dimensions of the artificial intelligence prediction model, plays an auxiliary role in modeling during the processes of generating the bandwidth status prediction series and the packet loss rate prediction series, improving the accuracy of the bandwidth status prediction series and the packet loss rate prediction series. Especially in high-concurrency live broadcast scenarios, the artificial intelligence prediction model can use the number of viewers feature to perceive the network congestion trend, thereby performing more accurate modeling of the packet loss risk or bandwidth fluctuations.
[0083] Among them, the fixed window time T and the predicted first preset time t are configurable preset times. The value range of the fixed window time T is 15 seconds to 60 seconds, and the value range of the predicted first preset time t is 5 seconds to 10 seconds. i is an integer between 1 and N, and N is the number of live streamers.
[0084] The artificial intelligence prediction model is trained through a collaborative live stream dataset, with the optimization goal of minimizing the error between the predicted value and the true value. During the training process, the mean squared error (MSE) of multivariable regression is used as the loss function to guide the model to continuously extract the time-dependent features in the input sequence, thereby improving the accuracy of predicting the bandwidth status, interaction delay, and packet loss rate within the predicted first preset time t.
[0085] The input of the artificial intelligence prediction model is a multi-dimensional time series of i live streamers at n consecutive time steps, and the multi-dimensional time series is as follows:
[0086]
[0087] Among them, i is an integer between 1 and N, N is the number of live streamers, and t is the predicted first preset time.
[0088] is the time series of the i-th live streamer within the predicted first preset time t, The expression of is:
[0089]
[0090] Among them, is the historical bandwidth time series of the i-th live streamer within the predicted first preset time t, is the historical interaction delay sequence of the i-th live streamer within the predicted first preset time t, is the historical packet loss rate time series of the i-th live streamer within the predicted first preset time t, is the historical jitter time series of the i-th live streamer within the predicted first preset time t, is the historical number of live streamers time series within the first preset time t, is the historical interaction frequency time series of the i-th live streamer within the predicted first preset time t, is the historical number of viewers time series of the i-th live streamer within the predicted first preset time t. i is an integer between 1 and N, and N is the number of live streamers.
[0091] The output of the artificial intelligence prediction model is the prediction sequence of i live streamers. Among them, the prediction sequence of the i-th live streamer is:
[0092]
[0093] Among them, is the bandwidth status prediction sequence for the i-th live streamer; is the packet loss rate prediction sequence for the i-th live streamer; is the interaction delay prediction sequence for the i-th live streamer, where i is an integer between 1 and N, and N is the number of live streamers.
[0094] The bandwidth status prediction sequence, packet loss rate prediction sequence, and interaction delay prediction sequence for the i-th live streamer are:
[0095]
[0096] Among them, is the bandwidth prediction value for the i-th live streamer at time t + 1; is the packet loss rate prediction value for the i-th live streamer at time t + 1; is the interaction delay prediction value for the i-th live streamer at time t + 1; t + n is the first preset time.
[0097] Step S103: Dynamically optimize the live data stream. Based on the bandwidth status prediction sequence, packet loss rate prediction sequence, and interaction delay prediction sequence, adaptively optimize the live data stream of each live streamer according to the live data stream adaptive strategy.
[0098] In this step, the live data stream adaptive strategy includes a dynamic bitrate adjustment strategy and a resolution and frame rate adjustment strategy;
[0099] The dynamic bitrate adjustment strategy is to calculate and adjust the dynamic bitrate of each live streamer in real time according to the bandwidth fluctuation prediction sequence, packet loss rate prediction sequence, and interaction delay prediction sequence. The calculation formula for the dynamic bitrate is:
[0100]
[0101] Among them, is the bandwidth fluctuation prediction sequence for the i-th live streamer, α is the weight coefficient of the bandwidth, β is the weight coefficient of the packet loss rate, is the packet loss rate prediction sequence for the i-th live streamer, is the weight coefficient of the interaction delay, is the interaction delay prediction sequence for the i-th live streamer.
[0102] In this embodiment, the weight coefficient can be preset, and the three weight coefficients need to meet the constraint condition to ensure the numerical stability and controllability of the dynamic bitrate calculation.
[0103] The resolution and frame rate adjustment strategy is to adjust the resolution and frame rate of each live streamer in real time according to the dynamic bitrate, including:
[0104] When the dynamic bit rate is greater than or equal to 6 Mbps, the resolution is 4K and the frame rate is 60 FPS;
[0105] When the dynamic bit rate is greater than or equal to 4 Mbps and less than 6 Mbps, the resolution is 2K and the frame rate is 60 FPS;
[0106] When the dynamic bit rate is greater than or equal to 2.5 Mbps and less than 4 Mbps, the resolution is 1080P and the frame rate is 60 FPS;
[0107] When the dynamic bit rate is greater than or equal to 1.5 Mbps and less than 2.5 Mbps, the resolution is 1080P and the frame rate is 30 FPS;
[0108] When the dynamic bit rate is greater than or equal to 1.0 Mbps and less than 1.5 Mbps, the resolution is 720P and the frame rate is 30 FPS;
[0109] When the dynamic bit rate is greater than or equal to 0.6 Mbps and less than 1.0 Mbps, the resolution is 480P and the frame rate is 30 FPS;
[0110] When the dynamic bit rate is greater than or equal to 0.4 Mbps and less than 0.6 Mbps, the resolution is 360P and the frame rate is 30 FPS.
[0111] Step S104: Cooperative live interaction optimization. Based on the interaction delay prediction sequence, construct an interaction experience function, globally evaluate the interaction delay of multiple hosts, and dynamically adjust the interaction between the host and the audience according to the cooperative live interaction adjustment strategy.
[0112] In this step, first, construct an interaction experience function according to the interaction delay prediction sequence. The interaction experience function is the trigger basis for the cooperative live interaction adjustment strategy, used to globally evaluate the interaction quality status, and optimize the interaction of all hosts according to the evaluation results to improve the interaction stability and consistency during the cooperative live broadcast.
[0113] When the interaction experience function is greater than the interaction optimization trigger threshold, execute the cooperative live interaction adjustment strategy.
[0114] The calculation formula of the interaction experience function is:
[0115]
[0116] Among them, is the penalty weight coefficient, used to adjust the sensitivity of the system to delay fluctuations; is the interaction delay prediction sequence of the i-th host; is the global interaction delay threshold.
[0117] In this embodiment, The larger the value, the more sensitive it is to the interaction delay differences among various live streamers, and it is more likely to trigger the collaborative live stream interaction adjustment strategy; The smaller the value, the more tolerant it is to the interaction delay differences among various live streamers, and the collaborative live stream interaction adjustment strategy is only triggered when the interaction delay differences are relatively large. The value can be preset, and the value range is between 0.5 - 2.0.
[0118] The formula for the global interaction delay threshold is:
[0119]
[0120] Among them, is the mean value of the interaction delay prediction sequences of all live streamers; is the adjustment coefficient, used to control the system's tolerance for delay jitter; is the standard deviation of the interaction delay prediction sequences of all live streamers.
[0121] In this embodiment, the larger the value, the more tolerant it is to the interaction delay differences among live streamers, and it is applicable to scenarios with a higher tolerance for interaction delay; the smaller the value, the more sensitive it is to the interaction delay differences, and it is more likely to trigger the collaborative live stream interaction adjustment strategy. The value can be preset, and the value range is between 0.1 - 1.0.
[0122] The interaction experience function QoE is used to measure the overall deviation degree of the current multi - live streamer interaction delay. The higher the value of the interaction experience function QoE, the greater the interaction delay among various live streamers and the worse the interaction experience. To globally determine whether the current live streamers need to initiate the collaborative live stream interaction adjustment strategy, an interaction optimization trigger threshold needs to be set to determine whether the interaction experience function QoE is within an acceptable range. The formula for the interaction optimization trigger threshold is:
[0123]
[0124] Among them, is the number of live streamers; is the maximum tolerable delay deviation tolerance value for a single live streamer.
[0125] In this embodiment, is used to determine whether the interaction delay of this live streamer exceeds the acceptable range, so as to decide whether this live streamer performs interaction optimization. The value can be preset, and the value range is between 100ms - 500ms.
[0126] Step S105: Multi-anchor video timing synchronization. Determine the synchronization reference time based on the video frame timestamps of all anchors, and perform corresponding video synchronization processing respectively according to the time difference between the video frame timestamp of each anchor and the synchronization reference time.
[0127] Figure 2 It is a flowchart of the collaborative live broadcast interaction adjustment strategy shown according to an exemplary embodiment of the present application. As Figure 2 shown, the collaborative live broadcast interaction adjustment strategy provided in this embodiment includes:
[0128] Step S201: Calculate the target interaction delay threshold for each anchor.
[0129] In this step, the target interaction delay threshold is a dynamic threshold, and the calculation formula of the target interaction delay threshold is:
[0130]
[0131] Wherein, is the mean value of the interaction delay prediction sequence of the i-th anchor; is the adjustment coefficient; is the standard deviation of the interaction delay prediction sequence of the i-th anchor.
[0132] In this embodiment, is used to calculate the target interaction delay threshold for each anchor to control the tolerance degree of the system to the interaction delay difference of this anchor. The value of can be preset, and its value range is between 0.1 and 1.0.
[0133] Step S202: Compare the interaction delay prediction sequence of each anchor with the corresponding target interaction delay threshold.
[0134] Step S203: When the interaction delay prediction sequence of the anchor is greater than the corresponding target interaction delay threshold, determine that this anchor is an interaction optimization object.
[0135] Step S204: Perform interaction optimization on the anchor determined to be an interaction optimization object. The interaction optimization includes reducing the bullet screen sending frequency, reducing the comment interaction frequency, simplifying the reward animation special effects, and combining and displaying the gift special effects.
[0136] Figure 3 It is a flowchart of the multi-anchor video timing synchronization shown according to an exemplary embodiment of the present application. As Figure 3 shown, the method provided in this embodiment includes:
[0137] Step S301: Generate a video frame timestamp sequence for each anchor based on the video frame timestamps collected within the second preset time.
[0138] In this step, it is necessary to count the video frame timestamp sequences of each anchor. For example, the frame rate of the i-th anchor is 30 FPS, and the video frame timestamp sequence within 1 second is as follows:
[0139]
[0140] Among them, is the timestamp of the first frame of the i-th anchor within this second; is the timestamp of the 30th frame of the i-th anchor within this second. If the second preset time is 2 seconds and the frame rate of the i-th anchor is 30 FPS, then the video frame timestamp sequence of this anchor contains 60 values of video frame timestamps.
[0141] Step S302: Determine the synchronization reference time by adopting the maximum value strategy according to the video frame timestamp sequences of all anchors.
[0142] In this step, the maximum value strategy is to select the largest video frame timestamp from the video frame timestamp sequences of all anchors as the synchronization reference time. Based on the synchronization reference time, perform frame interpolation operations or frame dropping operations on the video frames of other anchors. The maximum value strategy can ensure that the audio-visual content of all anchors does not play ahead, thus maintaining synchronization consistency. The calculation formula of the maximum value strategy is:
[0143]
[0144] where n is the number of anchors.
[0145] Step S303: Calculate the time difference between the video frame timestamp of each anchor and the synchronization reference time.
[0146] In this step, the calculation formula of the time difference is:
[0147]
[0148] Step S304: Perform frame interpolation operations or frame dropping operations on the video frames of each anchor respectively according to the time difference.
[0149] Including:
[0150] When the time difference is greater than 0, it means that the video frame of this anchor lags behind, and frame interpolation operations are performed on the video frame of this anchor;
[0151] When the time difference is less than 0, it means that the video frame of this anchor is ahead, and frame dropping operations are performed on the video frame of this anchor;
[0152] When the time difference is 0, it means that the video frame of this anchor is already aligned, and the video frame of this anchor is not processed.
[0153] In this step, the frame interpolation operation is a frame interpolation algorithm based on an artificial intelligence prediction model. By inputting adjacent video frames into the artificial intelligence prediction model, intermediate transition frames are predicted and generated. The frame interpolation algorithm uses the artificial intelligence prediction model to learn the motion features, image content changes, and temporal relationships between video frames, and generates one or more interpolated frames to compensate for the differences in the playback time of the host video stream.
[0154] In this embodiment, considering that the time difference between video frames of hosts during multi-host collaborative live streaming may exceed the single-frame interval, in order to achieve more precise video timing alignment, the artificial intelligence prediction model used for frame interpolation operation needs to support generating multiple intermediate transition frames. The input of this artificial intelligence prediction model is the adjacent frame images at any two time points, and according to the set time interpolation factor, any number of intermediate frames are output to fill the frame time difference.
[0155] The artificial intelligence prediction model used for frame interpolation operation can adopt a frame interpolation neural network structure guided by time control, such as the RIFE (Real-Time Intermediate Flow Estimation) framework. This type of model does not rely on complex optical flow modeling, and based on a lightweight convolutional structure, it completes the estimation of the inter-frame motion relationship and multi-frame fusion prediction, and has high efficiency and deployability.
[0156] The artificial intelligence prediction model used for frame interpolation operation can be trained using the Vimeo90K dataset. The training samples usually consist of three consecutive frame images, where the previous frame and the subsequent frame are used as inputs, and the middle frame is used as the supervision label. During the training process, this artificial intelligence prediction model aims to minimize the pixel error between the predicted generated intermediate frame and the real intermediate frame, and the loss function uses the L1 loss to optimize the frame interpolation accuracy and image restoration ability of the model.
[0157] The artificial intelligence prediction model used for frame interpolation operation is controlled in the way of inserting one frame per unit time interval. When the time difference of a certain host is detected then the frame interpolation operation is triggered. The input of the artificial intelligence prediction model is the current frame and the previous frame image of this host. According to the time difference of this host calculate the number of intermediate frames to be inserted, and set the corresponding time interpolation factor. Specifically, first preset a time step threshold Y, for example, the Y value is 33 milliseconds, and divide the time difference by the time step threshold Y, and the obtained quotient value is the number of intermediate frames to be inserted.
[0158] To insert multiple intermediate transition frames, time positions are equally divided between the current frame and the previous frame image, and multiple time interpolation factors are set. For example, when three frames need to be inserted, time interpolation factors of 0.25, 0.5, and 0.75 can be set respectively, indicating that the inserted frames are at the 1 / 4, 1 / 2, and 3 / 4 positions between the current frame and the previous frame image on the time axis.
[0159] The artificial intelligence prediction model for frame interpolation operation predicts and generates each intermediate transition frame in sequence based on the selected input frame image and interpolation factor, to fill the time difference of the anchor and achieve the timing alignment of multi-anchor video frames.
[0160] In this embodiment, when the time difference of a certain anchor is detected then the frame dropping operation is triggered. The frame dropping operation includes: first, find the frame with the timestamp closest to in the received frame sequence of the anchor as the target alignment frame; then discard all the leading frames later than this frame to ensure that the played frames are consistent with the synchronization reference time. A leading frame refers to a frame in the video frame sequence of an anchor whose timestamp is later than at the current synchronization moment. All such frames will be discarded to achieve the timing alignment of multi-anchor video frames, and only the frame closest to the timestamp is retained as the synchronized playback frame.
[0161] Figure 4 is a schematic structural diagram of a collaborative live interactive data processing system shown according to an exemplary embodiment of the present application. As Figure 4 shown, the collaborative live interactive data processing system 400 provided in this embodiment includes: a collaborative live data acquisition module 410, a data prediction module 420 based on an artificial intelligence prediction model, a live data flow dynamic optimization module 430, a collaborative live interactive optimization module 440, and a multi-anchor video timing synchronization module 450.
[0162] The collaborative live data acquisition module 410 acquires the anchor-side data and the platform-side data. The anchor-side data includes bandwidth, interaction delay, packet loss rate, jitter, and video frame timestamp, and the platform-side data includes the number of anchors, interaction frequency, and the number of audiences, and generates a multi-dimensional time series based on the anchor-side data and the platform-side data.
[0163] The data prediction module 420 based on an artificial intelligence prediction model, the artificial intelligence prediction model predicts the bandwidth status prediction sequence, packet loss rate prediction sequence, and interaction delay prediction sequence of each anchor within the first preset time based on the multi-dimensional time series.
[0164] The live data flow dynamic optimization module 430 dynamically optimizes the live data flow of each anchor according to the live data flow adaptive strategy based on the bandwidth status prediction sequence, packet loss rate prediction sequence, and interaction delay prediction sequence.
[0165] The collaborative live broadcast interaction optimization module 440 constructs an interaction experience function based on the interaction delay prediction sequence, globally evaluates the interaction delays of multiple hosts, and dynamically adjusts the interaction between the hosts and the audience according to the collaborative live broadcast interaction adjustment strategy.
[0166] The multi-host video time sequence synchronization module 450 determines the synchronization reference time based on the video frame timestamps of all hosts, and performs corresponding video synchronization processing respectively according to the time differences between the video frame timestamps of each host and the synchronization reference time.
[0167] Figure 5 It is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of the present application. As Figure 5 shown, an electronic device 500 provided in this embodiment includes: a processor 501 and a memory 502; wherein:
[0168] The memory 502 is used to store a computer program, and this memory can also be flash (flash memory).
[0169] The processor 501 is used to execute the execution instructions stored in the memory to implement each step in the above method. For specific reference, please refer to the relevant descriptions in the foregoing method embodiments.
[0170] Optionally, the memory 502 can be either independent or integrated with the processor 501.
[0171] When the memory 502 is a device independent of the processor 501, the electronic device 500 may further include:
[0172] A bus 503 for connecting the memory 502 and the processor 501.
[0173] This embodiment also provides a readable storage medium, in which a computer program is stored. When at least one processor of the electronic device executes this computer program, the electronic device executes the methods provided by the above various embodiments.
[0174] This embodiment also provides a program product, which includes a computer program, and this computer program is stored in a readable storage medium. At least one processor of the electronic device can read this computer program from the readable storage medium, and at least one processor executes this computer program to enable the electronic device to implement the methods provided by the above various embodiments.
[0175] Other embodiments of the present application will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0176] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A method for processing collaborative live interaction data, characterized in that, Including: Collaborative live broadcast data collection, collecting data on the host side and the platform side. The data on the host side includes bandwidth, interaction latency, packet loss rate, jitter, and video frame timestamps. The data on the platform side includes the number of hosts, interaction frequency, and the number of viewers. Generate a multi-dimensional time series based on the data on the host side and the platform side; Data prediction based on an artificial intelligence prediction model. The artificial intelligence prediction model is based on the multi-dimensional time series to predict the bandwidth status prediction sequence, packet loss rate prediction sequence, and interaction latency prediction sequence of each host within a first preset time; Dynamic optimization of the live broadcast data flow. Based on the bandwidth status prediction sequence, packet loss rate prediction sequence, and interaction latency prediction sequence, dynamically optimize the live broadcast data flow of each host according to the live broadcast data flow adaptive strategy; Collaborative live broadcast interaction optimization. Based on the interaction latency prediction sequence, construct an interaction experience function, globally evaluate the interaction latency of multiple hosts, and dynamically adjust the interaction between the host and the audience according to the collaborative live broadcast interaction adjustment strategy; Multi-host video time series synchronization. Determine the synchronization reference time based on the video frame timestamps of all hosts, and perform corresponding video synchronization processing according to the time difference between the video frame timestamp of each host and the synchronization reference time.
2. The collaborative live interaction data processing method according to claim 1, wherein, The live broadcast data flow adaptive strategy includes a dynamic bitrate adjustment strategy and a resolution and frame rate adjustment strategy; The dynamic bitrate adjustment strategy is to calculate and adjust the dynamic bitrate of each host in real time according to the bandwidth status prediction sequence, packet loss rate prediction sequence, and interaction latency prediction sequence; The resolution and frame rate adjustment strategy is to adjust the resolution and frame rate of each host in real time according to the dynamic bitrate, including: When the dynamic bitrate is greater than or equal to 6 Mbps, the resolution is 4K and the frame rate is 60 FPS; When the dynamic bitrate is greater than or equal to 4 Mbps and less than 6 Mbps, the resolution is 2K and the frame rate is 60 FPS; When the dynamic bitrate is greater than or equal to 2.5 Mbps and less than 4 Mbps, the resolution is 1080P and the frame rate is 60 FPS; When the dynamic bitrate is greater than or equal to 1.5 Mbps and less than 2.5 Mbps, the resolution is 1080P and the frame rate is 30 FPS; When the dynamic bitrate is greater than or equal to 1.0 Mbps and less than 1.5 Mbps, the resolution is 720P and the frame rate is 30 FPS; When the dynamic bitrate is greater than or equal to 0.6 Mbps and less than 1.0 Mbps, the resolution is 480P and the frame rate is 30 FPS; When the dynamic bitrate is greater than or equal to 0.4 Mbps and less than 0.6 Mbps, the resolution is 360P and the frame rate is 30 FPS.
3. The collaborative live interactive data processing method according to claim 1, wherein The interaction experience function is the trigger basis for the collaborative live broadcast interaction adjustment strategy; When the interaction experience function is greater than the interaction optimization trigger threshold, execute the collaborative live broadcast interaction adjustment strategy; The interaction optimization trigger threshold is used to determine whether the interaction experience function is within an acceptable range.
4. The collaborative live interactive data processing method according to claim 3, wherein The collaborative live broadcast interaction adjustment strategy includes: Calculate the target interaction latency threshold for each host; Compare the interaction delay prediction sequence of each said host with the corresponding said target interaction delay threshold; When the interaction delay prediction sequence of the host is greater than the corresponding said target interaction delay threshold, determine that the host is an object for interaction optimization; Perform interaction optimization on the hosts determined to be objects for interaction optimization, and the interaction optimization includes reducing the bullet screen sending frequency, reducing the comment interaction frequency, simplifying the reward animation special effects, and combining the display of gift special effects.
5. The collaborative live interactive data processing method according to claim 4, wherein The said target interaction delay threshold is a dynamic threshold.
6. The collaborative live interaction data processing method according to claim 1, wherein The multi-host video timing synchronization includes: Generate a video frame timestamp sequence for each host based on the video frame timestamps collected within a second preset time; Determine the synchronization reference time using the maximum value strategy according to the video frame timestamp sequences of all the hosts; Calculate the time difference between the video frame timestamp of each host and the synchronization reference time; According to the time difference, perform frame interpolation operations or frame dropping operations on the video frames of each host respectively, including: When the time difference is greater than 0, perform a frame interpolation operation on the video frames of the host; When the time difference is less than 0, perform a frame dropping operation on the video frames of the host; When the time difference is 0, the video frames of the host are not processed.
7. The collaborative live interactive data processing method according to claim 6, wherein, The frame interpolation operation is a frame interpolation algorithm based on an artificial intelligence prediction model, and by inputting adjacent video frames into the artificial intelligence prediction model, intermediate transition frames are predicted and generated.
8. A collaborative live interactive data processing system, characterized in that, Include: A collaborative live broadcast data collection module that collects host-side data and platform-side data. The host-side data includes bandwidth, interaction delay, packet loss rate, jitter, and video frame timestamps. The platform-side data includes the number of hosts, interaction frequency, and the number of viewers. Generate a multi-dimensional time series based on the host-side data and the platform-side data; A data prediction module based on an artificial intelligence prediction model. The artificial intelligence prediction model is based on the multi-dimensional time series and predicts the bandwidth status prediction sequence, packet loss rate prediction sequence, and interaction delay prediction sequence of each host within a first preset time; A live broadcast data flow dynamic optimization module that dynamically optimizes the live broadcast data flow of each host according to the live broadcast data flow adaptive strategy based on the bandwidth status prediction sequence, packet loss rate prediction sequence, and interaction delay prediction sequence; A collaborative live broadcast interaction optimization module that constructs an interaction experience function based on the interaction delay prediction sequence, globally evaluates the multi-host interaction delay, and dynamically adjusts the interaction between the host and the audience according to the collaborative live broadcast interaction adjustment strategy; A multi-host video timing synchronization module that determines the synchronization reference time based on the video frame timestamps of all the hosts, and performs corresponding video synchronization processing according to the time difference between the video frame timestamp of each host and the synchronization reference time.
9. An electronic device, characterized in that, Include: A processor; And, A memory for storing the executable instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1 to 7 by executing the executable instructions.
10. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by the processor, they are used to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Mobile terminal audio and video synchronization method and device, equipment and storage medium
CN113225598A
Large-scale digital live broadcast method and system with multiple cooperative devices
CN118678115A