WebRTC audio and video stream parameter dynamic adjustment method and device and computer equipment
Through the comprehensive scoring and adaptive model, the problem of difficulty in automatically adjusting the audio and video streaming parameters in the existing technology is solved, and more efficient audio and video streaming transmission and better user experience are achieved.
Patent Information
- Application Number
- CN202411984502.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
Smart Images

Figure CN120017644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a parameter adjustment method, and more specifically to a WebRTC audio and video stream parameter dynamic adjustment method, device and computer equipment. Background Art
[0002] WebRTC (Web Real-Time Communication) is an open source project and technical specification that aims to enable real-time audio and video communication between browsers. It provides powerful real-time communication capabilities for web developers through a simple JavaScript API and is widely used in video conferencing, online education, social media and other fields. The main features of WebRTC include low latency, high-quality audio and video transmission, and peer-to-peer communication.
[0003] Existing WebRTC implementations often rely on fixed audio and video streaming parameters, such as resolution and frame rate. This makes it difficult to automatically adjust when network conditions change, resulting in reduced audio and video quality. The encoder optimization used in this process relies mainly on fixed parameter settings, and fails to fully utilize real-time monitoring data for dynamic adjustments. Existing products often cannot be flexibly adjusted according to the specific needs of users. User settings are usually static, lack intelligent analysis and adaptive capabilities, and cannot respond to changes in user preferences in a timely manner.
[0004] Therefore, it is necessary to design a new method to dynamically adjust parameters to adapt to the dynamically changing network environment and improve transmission efficiency and quality. Summary of the invention
[0005] The purpose of the present invention is to overcome the defects of the prior art and provide a method, device and computer equipment for dynamically adjusting WebRTC audio and video stream parameters.
[0006] To achieve the above object, the present invention adopts the following technical solution: a method for dynamically adjusting parameters of WebRTC audio and video streams, comprising:
[0007] Get WebRTC audio and video streams and current network status;
[0008] Performing a comprehensive score calculation on the WebRTC audio and video stream to obtain a calculation result;
[0009] When the calculation result is less than a set threshold, the current network state is input into the adaptive model to obtain an adjustment strategy;
[0010] Dynamically adjust the buffer size of audio and video data according to the adjustment strategy;
[0011] Adjust the GOP size, QP value, and bitrate control based on real-time monitoring of WebRTC audio and video streams and user preferences.
[0012] The further technical solution is: the obtaining of WebRTC audio and video streams and current network status includes:
[0013] Create an RTCPeerConnection instance to obtain the WebRTC audio and video stream, and periodically call the getStats() method to obtain real-time statistics of the WebRTC audio and video stream;
[0014] Key indicators are extracted from the real-time statistical data to obtain the current network status, wherein the key indicators include delay, packet loss rate, bandwidth, frame rate, resolution, and encoding format.
[0015] The further technical solution is: the comprehensive scoring calculation of the WebRTC audio and video stream to obtain the calculation result includes:
[0016] Obtain user feedback information on the WebRTC audio and video stream;
[0017] A comprehensive score is calculated according to the feedback information and the current network status using an indicator weighted average method to obtain a calculation result.
[0018] Its further technical solution is: the adaptive model is obtained by combining historical data and real-time data to construct a situational model trained by a random forest algorithm, wherein the historical data and real-time data include network status at different time periods and user feedback information.
[0019] The further technical solution is: the adaptive model is obtained by combining historical data and real-time data with a situational model constructed by random forest algorithm training, including:
[0020] Get historical data as well as real-time data;
[0021] Preprocessing the historical data and the real-time data to obtain a preprocessing result;
[0022] Performing feature selection on the preprocessing result to obtain a selection result;
[0023] Dividing the network state into different scenarios according to the selection result, and defining corresponding labels and adjustment strategies to obtain a scenario model;
[0024] The situational model is trained using a random forest algorithm to obtain an adaptive model.
[0025] A further technical solution is: dynamically adjusting the buffer size of audio and video data according to the adjustment strategy includes:
[0026] According to the adjustment strategy combined with the time series analysis method, the future network status change is predicted to dynamically adjust the buffer size of the audio and video data.
[0027] The further technical solution is: adjusting the GOP size, QP value and bit rate control according to the real-time monitored WebRTC audio and video stream and user preferences, including:
[0028] Determine optimization priorities based on user preferences;
[0029] Combined with real-time monitoring of WebRTC audio and video streams, the GOP size, QP value and bitrate control are adjusted according to the optimization level.
[0030] The present invention also provides a device for dynamically adjusting parameters of WebRTC audio and video streams, including:
[0031] The acquisition unit is used to obtain the WebRTC audio and video streams and the current network status;
[0032] A calculation unit, configured to perform a comprehensive score calculation on the WebRTC audio and video stream to obtain a calculation result;
[0033] An adjustment strategy generating unit, configured to input the current network state into an adaptive model to obtain an adjustment strategy when the calculation result is less than a set threshold;
[0034] A first adjustment unit, configured to dynamically adjust the buffer size of the audio and video data according to the adjustment strategy;
[0035] The second adjustment unit is used to adjust the GOP size, QP value and bit rate control according to the real-time monitored WebRTC audio and video streams and user preferences.
[0036] The present invention further provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor implements the above method when executing the computer program.
[0037] The present invention also provides a storage medium, wherein the storage medium stores a computer program, and the computer program implements the above method when executed by a processor.
[0038] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention obtains real-time data by monitoring the quality of audio and video streams and the current network status; comprehensively scores the quality of audio and video streams to evaluate their current transmission effect and determine whether they meet the set quality standards; when the score is lower than the set threshold, the network status is input into the adaptive model to generate an appropriate adjustment strategy; the buffer size of audio and video data is dynamically adjusted according to the adjustment strategy to optimize transmission performance and reduce jamming; and the GOP size, QP value and bit rate control are adjusted in combination with real-time stream data and user preferences, thereby optimizing the quality and transmission efficiency of the video stream.
[0039] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0041] Figure 1 A schematic diagram of a process for dynamically adjusting WebRTC audio and video stream parameters provided by an embodiment of the present invention;
[0042] Figure 2 A schematic diagram of a sub-process of a method for dynamically adjusting WebRTC audio and video stream parameters provided by an embodiment of the present invention;
[0043] Figure 3 A schematic diagram of a sub-process of a method for dynamically adjusting WebRTC audio and video stream parameters provided by an embodiment of the present invention;
[0044] Figure 4 A schematic diagram of a sub-process of a method for dynamically adjusting WebRTC audio and video stream parameters provided by an embodiment of the present invention;
[0045] Figure 5 A schematic diagram of a sub-process of a method for dynamically adjusting WebRTC audio and video stream parameters provided by an embodiment of the present invention;
[0046] Figure 6 A schematic block diagram of a device for dynamically adjusting WebRTC audio and video stream parameters provided by an embodiment of the present invention;
[0047] Figure 7 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0048] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0049] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0050] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0051] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0052] See also Figure 1 , Figure 1 A schematic flow chart of a method for dynamically adjusting parameters of WebRTC audio and video streams provided in an embodiment of the present invention. The method for dynamically adjusting parameters of WebRTC audio and video streams is applied to a server. The server interacts with the terminal to perform data exchange, first by monitoring the status of the WebRTC audio and video streams and key network indicators in real time, calculating a comprehensive score and obtaining user feedback. Next, an adaptive model (such as a random forest algorithm) is used to train a scenario model based on historical and real-time data to generate a dynamic adjustment strategy. According to the adjustment strategy, the buffer size of the audio and video data is dynamically adjusted in combination with time series analysis to predict future network changes. Finally, according to user preferences and real-time monitoring results, the GOP size, QP value, and bit rate control are adjusted to optimize the video stream quality.
[0053] Figure 1 Schematic diagram of the process of the method for dynamically adjusting WebRTC audio and video stream parameters provided by an embodiment of the present invention. Figure 1 As shown, the method includes the following steps S110 to S150.
[0054] S110, obtaining WebRTC audio and video streams and current network status.
[0055] In this embodiment, the current network status mainly refers to performance indicators such as packet loss rate and delay during network communication, which directly affect the quality of audio and video calls.
[0056] In one embodiment, see Figure 2 , the above-mentioned step S110 may include steps S111 to S112.
[0057] S111. Create an RTCPeerConnection instance to obtain the WebRTC audio and video stream, and periodically call the getStats() method to obtain real-time statistical data of the WebRTC audio and video stream.
[0058] In this embodiment, when the application starts, an RTCPeerConnection instance is created by calling new RTCPeerConnection(config), where config may include information such as ICE server configuration to help establish a P2P connection.
[0059] Add audio and video tracks to the RTCPeerConnection instance using the addTrack(track, stream) method. These tracks come from local media streams (for example, obtained using getUserMedia).
[0060] Complete SDP negotiation (Offer / Answer exchange) and ICE candidate exchange to establish a connection between the two parties.
[0061] Set a timer (for example, using setInterval) to call the getStats() method of the RTCPeerConnection instance at regular intervals (for example, once a second). This method returns a Promise that resolves to provide statistics about the current connection.
[0062] S112. Extract key indicators from the real-time statistical data to obtain the current network status, wherein the key indicators include delay, packet loss rate, bandwidth, frame rate, resolution, and encoding format.
[0063] In this embodiment, when the Promise returned by the getStats() method is resolved, an object containing various types of statistical information will be returned, including but not limited to sender, receiver, transport, ICE candidate, etc.
[0064] From the returned statistics, we extract several indicators that are critical to evaluating the network status:
[0065] Latency (RTT): This can usually be found by looking at transport-level statistics.
[0066] Packet loss rate: can be obtained from the statistics of the sender or receiver.
[0067] Bandwidth: includes the estimated available bandwidth for both upstream and downstream, which can also be obtained from the statistics of the sender or receiver.
[0068] Frame rate: The frame rate of the video stream, obtained from the statistics of the sender or receiver.
[0069] Resolution: The resolution of the video stream, which can also be obtained from the statistics of the sender or receiver.
[0070] Coding format: The coding format used by the audio and video streams, obtained from the statistics of the sender or receiver.
[0071] These key metrics can be used to dynamically adjust the parameters of the WebRTC session, such as reducing the resolution or frame rate to reduce bandwidth usage, or notifying the user that the current network conditions are poor and may affect the call quality.
[0072] S120: Perform comprehensive score calculation on the WebRTC audio and video stream to obtain a calculation result.
[0073] In this embodiment, the calculation result refers to the comprehensive score of the WebRTC audio and video stream.
[0074] In one embodiment, see Figure 3 , the above-mentioned step S120 may include steps S121 to S122.
[0075] S121. Obtain user feedback information on the WebRTC audio and video stream.
[0076] In this embodiment, a feedback form is provided on the user interface to allow users to evaluate the audio and video experience. The form may include but is not limited to ratings for fluency, picture clarity, audio quality, etc., usually in the form of star ratings (such as 1 to 5 stars).
[0077] After the user completes the call, prompt the user to fill out a feedback form. The collected feedback information should include specific values for each rating, such as 4 stars for fluency and 5 stars for picture clarity.
[0078] S122: Calculate a comprehensive score using an indicator weighted average method according to the feedback information and the current network status to obtain a calculation result.
[0079] In this embodiment, a weight is assigned to each indicator according to the degree of influence of each indicator on the quality of the audio and video stream. For example: RTT (delay) is 0.4; packet loss rate is 0.4; user satisfaction is 0.2; the selection of weights should be based on experience or experimental data to ensure that the impact of each indicator on the overall quality can be accurately reflected.
[0080] Since different indicators have different units and magnitudes, it is necessary to standardize all indicator values to the same range (for example, between 0 and 1). For example: Packet loss rate: 1-packet loss rate; user satisfaction directly uses user ratings (assuming 1 star is 0 and 5 stars is 1).
[0081] The weighted average formula is used to calculate the comprehensive score. Assume that there are n indicators, each indicator x i The weight is w i , then the comprehensive score S can be expressed as:
[0082] For example: suppose there is the following data:
[0083] RTT = 100ms (weight 0.4);
[0084] Packet loss rate = 2% (weight 0.4);
[0085] User satisfaction = 8 / 10 (converted to 0.8, weight 0.2);
[0086] The comprehensive score S is calculated as follows:
[0087] The system sets up a scheduled task (for example, every 30 minutes) to collect user feedback, and combines the new feedback data with real-time network monitoring data to update the comprehensive score. This ensures that the decision is based on the overall quality of the current audio and video stream.
[0088] According to the system design, a comprehensive score threshold is set. When the comprehensive score is lower than the threshold, the adaptive decision algorithm is triggered. Once the trigger mechanism is started, the adaptive decision algorithm will dynamically adjust the encoder parameters, buffer size, etc. according to the current network status and user preferences to improve the quality of audio and video streams.
[0089] S130: When the calculation result is less than a set threshold, the current network state is input into an adaptive model to obtain an adjustment strategy.
[0090] In this embodiment, the adjustment strategy refers to the result obtained by the adaptive model by combining historical data and real-time data and training the context model constructed by the random forest algorithm. The adjustment strategy is intended to optimize the transmission quality of the WebRTC audio and video streams and ensure the best user experience under different network conditions.
[0091] The adaptive model is obtained by combining historical data and real-time data to build a situational model trained by a random forest algorithm, wherein the historical data and real-time data include network status at different time periods and user feedback information.
[0092] In one embodiment, see Figure 4 The above-mentioned adaptive model is obtained by combining historical data and real-time data to train a situational model constructed by random forest algorithm, including steps S131 to S135.
[0093] S131. Obtain historical data and real-time data.
[0094] In this embodiment, network status data over a period of time is collected, including but not limited to RTT, packet loss rate, bandwidth, etc. User feedback data is collected, such as scores of audio and video fluency, picture clarity, audio quality, etc., to obtain historical data.
[0095] Monitor the current network status in real time, including RTT, packet loss rate, bandwidth, etc.; collect user feedback information in real time, such as scores obtained through user interface forms, to obtain real-time data.
[0096] S132: preprocess the historical data and the real-time data to obtain a preprocessing result.
[0097] In this embodiment, outliers and missing values are removed to ensure the accuracy and completeness of the data. For example, if a data point is significantly out of the normal range, it can be considered an outlier and removed.
[0098] Normalize all feature values to the same range (e.g., between 0 and 1) to ensure comparability between different features.
[0099] S133: Perform feature selection on the preprocessing result to obtain a selection result.
[0100] In this embodiment, features that have a greater impact on the quality of audio and video streams are selected, such as RTT, packet loss rate, bandwidth, and user feedback.
[0101] You can use feature importance analysis methods such as feature importance of random forests to determine which features are most effective in contributing to the model's predictions.
[0102] S134. Divide the network status into different situations according to the selection result, and define corresponding labels and adjustment strategies to obtain a situation model.
[0103] In this embodiment, based on historical data and real-time data, the network status is divided into multiple scenario categories, such as: normal status; high packet loss status; high latency status; insufficient bandwidth status.
[0104] Define labels for each scenario category and record the corresponding adjustment strategies. For example: in the case of high packet loss, reduce the resolution and frame rate; in the case of high latency, increase the buffer size; in the case of insufficient bandwidth, reduce the bit rate.
[0105] S135. Use a random forest algorithm to train the situation model to obtain an adaptive model.
[0106] In this embodiment, a labeled data set is used for model training, the input feature is the network state, and the output is the corresponding adjustment strategy. The training is performed using the random forest algorithm, which is an integrated learning method that can effectively process high-dimensional data and prevent overfitting. The model parameters are optimized through the back propagation algorithm to improve the prediction accuracy.
[0107] When the calculation result (comprehensive score) is lower than the set threshold, the current network status is input into the trained adaptive model. The adaptive model predicts the most suitable adjustment strategy based on the current network status and applies it to the transmission parameter adjustment of WebRTC audio and video streams.
[0108] The above step S110 provides real-time monitoring data as an important input to the situation model. These real-time data help identify the current network status and map it to predefined situation categories.
[0109] The situation model will continue to receive data updates from step S110 during operation. When network conditions change, the situation model will be triggered to re-evaluate the current situation and select the best adjustment strategy based on the latest data.
[0110] The situational model not only relies on historical data, but also combines real-time information, allowing the system to respond quickly to dynamic changes. This combination ensures that the decision-making process is highly flexible and accurate, thereby improving the quality and stability of WebRTC audio and video streaming.
[0111] Through the above steps, the situational model can effectively learn and adapt to the optimal parameter settings under different networks, providing strong support for dynamically adjusting WebRTC audio and video streams.
[0112] S140. Dynamically adjust the buffer size of the audio and video data according to the adjustment strategy.
[0113] In this embodiment, the adjustment strategy is combined with a time series analysis method to predict future changes in network status, so as to dynamically adjust the buffer size of audio and video data.
[0114] Specifically, receive real-time network status data, including but not limited to RTT (round trip time), packet loss rate, and bandwidth information. These data are used to evaluate the current network status in real time and provide a basis for adjusting the buffer size. Monitor changes in network status. When an increase in network latency or an increase in packet loss rate is detected, the system will automatically increase the buffer size for audio and video data. For example, if the RTT exceeds the set threshold, the system will increase the buffer capacity to maintain smooth playback during network fluctuations. Conversely, when the network condition is good, the buffer size can be appropriately reduced to reduce resource usage.
[0115] Of course, you can also use prediction to plan ahead; collect and analyze past network fluctuation data, including RTT, packet loss rate, bandwidth changes and other information. Identify common patterns and trends to provide a basis for predicting future network status. Use time series analysis methods to predict future network status. Specifically, based on historical latency and bandwidth data, predict network performance in the next few seconds. Time series analysis methods can use algorithms such as ARIMA (autoregressive integrated moving average model) and LSTM (long short-term memory network), which can capture long-term and short-term trends in time series data. Adjust the buffer size in advance based on the prediction results. For example, if high latency is predicted, increase the buffer size in advance to cope with possible degradation in streaming quality. This predictive adjustment helps to prepare before actual network fluctuations occur, thereby reducing playback jams and image quality degradation.
[0116] The specific process of dynamic adjustment is as follows:
[0117] An initial buffer size is set based on the current network status and historical data. For example, the initial buffer size can be set to a default value or preset based on the record of the last network status.
[0118] Receive real-time data regularly and dynamically adjust the buffer size based on the current network status and prediction results.
[0119] The adjustment strategy may include but is not limited to: Increase buffer size: When the RTT exceeds the threshold or the packet loss rate increases, increase the buffer capacity. Reduce buffer size: When the network condition is good and the prediction results show that the future network state is stable, appropriately reduce the buffer size.
[0120] When network conditions are good, unnecessary data caching is reduced, thereby reducing storage requirements and processing burden, and improving system efficiency. Dynamic adjustment of buffer size not only improves playback quality, but also optimizes the use of system resources.
[0121] By dynamically adjusting the buffer size, the playback freeze and image quality degradation caused by network fluctuations are effectively reduced, thereby improving the user experience. With the help of prediction algorithms, potential network problems can be identified in advance and preventive adjustments can be made, making the system more flexible and reliable when facing emergencies. Dynamic adjustment of the buffer not only improves playback quality, but also optimizes the use of system resources. In a good network environment, unnecessary data caching is reduced, thereby reducing storage requirements and processing burdens, and improving system efficiency. Through dynamic adjustment and prediction, stable and efficient support is provided for the outflow of WebRTC audio and video streams.
[0122] S150: Adjust the GOP size, QP value and bit rate control according to the real-time monitored WebRTC audio and video streams and user preferences.
[0123] In this embodiment, GOP (Group of Pictures) is an important concept in video encoding, which refers to a sequence consisting of a group of continuous image frames, which include at least one I frame (key frame) and subsequent P frames (forward prediction frames) and B frames (bidirectional prediction frames). The length of GOP defines the distance between two I frames, which directly affects the compression efficiency and decoding performance of the video.
[0124] I frame (key frame): I frame is a completely self-contained frame that provides complete picture information and can be decoded without reference to other frames. The first frame of each GOP must be an I frame.
[0125] P frame (forward prediction frame): P frame is predictively encoded by referring to one or more previous I frames or P frames, thereby reducing redundant information and improving compression efficiency.
[0126] B frame (bidirectional prediction frame): B frame is predicted and encoded by referring to one or more previous and subsequent I frames or P frames, further improving compression efficiency.
[0127] The setting of GOP has a significant impact on video quality, coding efficiency and random access capability. A longer GOP can improve compression efficiency but increase decoding delay; a shorter GOP can reduce delay but reduce compression efficiency.
[0128] QP (Quantization Parameter) is a key parameter in video encoding, which is used to control the degree of image compression. The QP value determines the degree of detail retention of each macroblock, thus affecting the visual quality and bit rate of the final video.
[0129] Low QP value: means less compression, higher image quality, but larger file size.
[0130] A high QP value means a higher degree of compression, lower image quality, but a smaller file size.
[0131] During the encoding process, by dynamically adjusting the QP value, the video quality can be optimized according to the current network conditions and user needs.
[0132] CBR (Constant Bit Rate) is a video encoding mode in which the encoder maintains a fixed bit rate throughout the video stream. This means that the amount of data transmitted per second is constant, regardless of the complexity of the video content.
[0133] It is suitable for application scenarios with strict bandwidth requirements, such as live broadcast and real-time communication, and can ensure stable data flow and reduce delays and buffering.
[0134] In situations where the scene changes significantly, the CBR mode may cause fluctuations in image quality because it cannot dynamically adjust the bitrate to adapt to the complexity of the content.
[0135] VBR (Variable Bit Rate) is a video encoding mode in which the encoder dynamically adjusts the bit rate according to the complexity of the video content. For static scenes, the encoder can use a lower bit rate; in complex or dynamic scenes, the bit rate will be increased to maintain image quality.
[0136] It is more common in many applications, such as video on demand and archiving, as it can optimize storage space usage while maintaining visual quality.
[0137] Since the bit rate is not fixed, VBR may cause uncertainty in bandwidth usage, which may not be suitable for some real-time applications.
[0138] By properly setting the GOP and QP values and selecting the appropriate bit rate control mode (CBR or VBR), you can effectively optimize the performance of video encoding, improve video quality, reduce latency, and adapt to different network conditions and user needs. The dynamic adjustment of these parameters is an important means in modern video encoding technology and can significantly improve user experience.
[0139] In one embodiment, see Figure 5, the above-mentioned step S150 may include steps S151 to S152.
[0140] S151, determining an optimization priority level according to user preference;
[0141] S152, combining the real-time monitored WebRTC audio and video streams, adjusting the GOP size, QP value and bit rate control according to the optimization level.
[0142] In this embodiment, an intuitive user interface is provided to enable users to select options such as prioritizing image quality, smoothness, or low latency. These options will affect the adjustment strategy of encoder parameters. By analyzing the user's behavior during use (such as viewing time, selected settings, playback quality, etc.), data is collected to identify user habits. Based on the results of behavioral analysis, the system can automatically adjust the audio and video stream parameters to adapt to user habits.
[0143] User preferences include the following:
[0144] Prioritize image quality: When this option is selected, the system will optimize the encoder parameters to provide the highest quality video picture, even if this may sacrifice some smoothness or increase latency.
[0145] Prioritize smoothness: When this option is selected, the system will optimize encoder parameters to reduce freezes and buffering time, ensuring smoother video playback even if the image quality is reduced.
[0146] Low latency priority: When this option is selected, the system will optimize the encoder parameters to reduce transmission delay, which is suitable for real-time interactive scenarios such as live broadcasts or video calls, even if the image quality and smoothness are reduced.
[0147] Record the duration of each viewing session and analyze user preferences for different settings. Record the user's preferred settings to understand the user's long-term preferences. Monitor indicators such as the number of freezes and buffering time during playback to evaluate the user experience under different settings.
[0148] Initialize the encoder parameters based on the user's first selected preference settings. During user use, dynamically adjust the encoder parameters based on the collected data to meet the user's real-time needs. Based on the results of user behavior analysis, the system can intelligently recommend the optimal setting combination to further improve the user experience.
[0149] By providing a variety of user preference settings, the system can better meet the needs of different users and improve user satisfaction. The function of dynamically adjusting parameters enables the system to flexibly respond to different network environments and user behaviors and provide more personalized services. Users can choose priorities according to their needs to ensure the best viewing experience in specific scenarios. Through intelligent analysis and automatic adjustment, the system reduces the tedious operation of manual adjustment by users and improves the convenience and comfort of use. By reasonably allocating network bandwidth and computing resources, the system can maximize resource utilization while ensuring user experience. Dynamic adjustment of parameters can also reduce unnecessary resource waste and improve the overall performance and stability of the system.
[0150] In addition, the system reads the user's preferences to determine the priority of optimization. For example, if the user selects "Prioritize smoothness", the system will prioritize reducing lag and buffering time.
[0151] Real-time monitoring data includes:
[0152] Network status: including RTT (Round-Trip Time), packet loss rate, bandwidth and other indicators.
[0153] Playback quality: includes indicators such as the number of freezes, buffering time, and playback smoothness.
[0154] When the network is good, increase the GOP size to improve encoding efficiency. When the network is bad, reduce the GOP size to reduce latency. When bandwidth is sufficient, use a lower QP value to improve image quality. When bandwidth is insufficient, increase the QP value to reduce the bitrate.
[0155] Real-time bandwidth conditions: Select the appropriate bit rate control mode (such as CBR or VBR) to ensure the stability of the video stream.
[0156] When the network is in good condition, the user selects "Prioritize image quality": Increase the GOP size to improve encoding efficiency. Use a lower QP value to improve image quality. Select CBR mode to ensure stable bitrate output.
[0157] When the network condition is poor, the user selects "Prioritize Fluency": Reduce GOP size and reduce latency. Increase QP value and reduce bitrate. Select VBR mode to dynamically adjust bitrate according to network bandwidth.
[0158] When the network condition is poor, the user selects "Prioritize low latency": reducing the GOP size and reducing latency.
[0159] Increase the QP value and reduce the bit rate. Select VBR mode to ensure the lowest latency.
[0160] If a high RTT and increased packet loss rate are detected, the system will automatically adjust the encoder parameters: reduce the GOP size to reduce latency, increase the QP value, and reduce the bit rate. Select VBR mode to ensure smooth video playback.
[0161] The above method significantly improves the network adaptability of the system by monitoring the network status in real time and dynamically adjusting the audio and video stream parameters (such as resolution and frame rate). This method can automatically reduce the stream quality when the network condition is poor, thereby maintaining the continuity of transmission and reducing the jamming phenomenon. A personalized setting interface is provided, allowing users to choose to prioritize image quality, fluency or low latency. The system not only allows manual settings, but also predicts user needs through intelligent analysis and makes automatic adjustments, thereby improving the personalized experience. By dynamically adjusting the encoder parameters (such as GOP size, QP value), combined with content-aware technology, the encoding strategy is optimized according to the characteristics of the video content. This method not only improves the encoding efficiency, but also optimizes the resource utilization, so that the system can maintain good transmission quality under various network conditions. The introduction of an adaptive decision-making algorithm based on deep learning continuously optimizes the situational model by analyzing historical data and real-time monitoring results, so that the system can self-learn and adapt to the optimal parameter settings under different network environments. This intelligent decision-making mechanism significantly improves the flexibility and adaptability of the system.
[0162] The method of this embodiment significantly improves the system's network adaptability, user experience, coding efficiency, and intelligent decision-making capabilities by real-time monitoring of network status, providing personalized settings, dynamically optimizing encoder parameters, and introducing a self-learning mechanism. These improvements enable this patent to perform well in a variety of application scenarios, especially when network conditions are complex and changeable.
[0163] The above-mentioned WebRTC audio and video stream parameter dynamic adjustment method obtains real-time data by monitoring the quality of audio and video streams and the current network status; comprehensively scores the quality of audio and video streams to evaluate their current transmission effect and determine whether they meet the set quality standards; when the score is lower than the set threshold, the network status is input into the adaptive model to generate an appropriate adjustment strategy; the buffer size of audio and video data is dynamically adjusted according to the adjustment strategy to optimize transmission performance and reduce jamming; and the GOP size, QP value and bit rate control are adjusted in combination with real-time stream data and user preferences, thereby optimizing the quality and transmission efficiency of the video stream.
[0164] Figure 6 is a schematic block diagram of a WebRTC audio and video stream parameter dynamic adjustment device 300 provided in an embodiment of the present invention. Figure 6As shown, corresponding to the above WebRTC audio and video stream parameter dynamic adjustment method, the present invention also provides a WebRTC audio and video stream parameter dynamic adjustment device 300. The WebRTC audio and video stream parameter dynamic adjustment device 300 includes a unit for executing the above WebRTC audio and video stream parameter dynamic adjustment method, and the device can be configured in a server. Specifically, please refer to Figure 6 The WebRTC audio and video stream parameter dynamic adjustment device 300 includes an acquisition unit 301, a calculation unit 302, an adjustment strategy generation unit 303, a first adjustment unit 304 and a second adjustment unit 305.
[0165] The acquisition unit 301 is used to obtain the WebRTC audio and video stream and the current network status; the calculation unit 302 is used to perform a comprehensive score calculation on the WebRTC audio and video stream to obtain a calculation result; the adjustment strategy generation unit 303 is used to input the current network status into the adaptive model when the calculation result is less than a set threshold to obtain an adjustment strategy; the first adjustment unit 304 is used to dynamically adjust the buffer size of the audio and video data according to the adjustment strategy; the second adjustment unit 305 is used to adjust the GOP size, QP value and bit rate control according to the real-time monitored WebRTC audio and video stream and user preferences.
[0166] In one embodiment, the acquisition unit 301 includes:
[0167] Create a subunit to create an RTCPeerConnection instance to obtain the WebRTC audio and video stream, and periodically call the getStats() method to obtain real-time statistics of the WebRTC audio and video stream;
[0168] The extraction subunit is used to extract key indicators from the real-time statistical data to obtain the current network status, wherein the key indicators include delay, packet loss rate, bandwidth, frame rate, resolution, and encoding format.
[0169] In one embodiment, the calculation unit 302 includes:
[0170] An information acquisition subunit, used to obtain user feedback information on the WebRTC audio and video stream;
[0171] The score calculation subunit is used to calculate the comprehensive score by using the indicator weighted average method according to the feedback information and the current network status to obtain a calculation result.
[0172] In one embodiment, the device further includes a model training unit, and the model training unit includes:
[0173] A data acquisition subunit, used to acquire historical data and real-time data;
[0174] A preprocessing subunit, used for preprocessing the historical data and the real-time data to obtain a preprocessing result;
[0175] A feature selection subunit, used for performing feature selection on the preprocessing result to obtain a selection result;
[0176] A model building subunit, used to divide the network state into different scenarios according to the selection result, and define corresponding labels and adjustment strategies to obtain a scenario model;
[0177] The training subunit is used to train the scenario model using a random forest algorithm to obtain an adaptive model.
[0178] In one embodiment, the first adjustment unit 304 is used to predict future changes in network status according to the adjustment strategy in combination with a time series analysis method, so as to dynamically adjust the buffer size of the audio and video data.
[0179] In one embodiment, the second adjustment unit 305 includes:
[0180] A level determination subunit, used to determine the optimization priority level according to user preferences;
[0181] The adjustment subunit is used to adjust the GOP size, QP value and bit rate control according to the optimization level in combination with the real-time monitored WebRTC audio and video streams.
[0182] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned WebRTC audio and video stream parameter dynamic adjustment device 300 and each unit can refer to the corresponding description in the aforementioned method embodiment, and for the convenience and brevity of description, it will not be repeated here.
[0183] The WebRTC audio and video stream parameter dynamic adjustment device 300 can be implemented in the form of a computer program. The computer program can be implemented in the form of a computer program. Figure 7 Runs on the computer device shown.
[0184] See also Figure 7 , Figure 7 5 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 may be a server, wherein the server may be an independent server or a server cluster composed of multiple servers.
[0185] See also Figure 7 The computer device 500 includes a processor 502 , a memory and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .
[0186] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, and when the program instructions are executed, the processor 502 can execute a method for dynamically adjusting parameters of a WebRTC audio and video stream.
[0187] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500 .
[0188] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a method for dynamically adjusting WebRTC audio and video stream parameters.
[0189] The network interface 505 is used to communicate with other devices over the network. Figure 7 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0190] The processor 502 is used to run the computer program 5032 stored in the memory to implement the following steps:
[0191] Acquire WebRTC audio and video streams and current network status; perform comprehensive scoring calculation on the WebRTC audio and video streams to obtain a calculation result; when the calculation result is less than a set threshold, input the current network status into an adaptive model to obtain an adjustment strategy; dynamically adjust the buffer size of audio and video data according to the adjustment strategy; adjust the GOP size, QP value and bit rate control according to the real-time monitored WebRTC audio and video streams and user preferences.
[0192] The adaptive model is obtained by combining historical data and real-time data to build a situational model trained by a random forest algorithm, wherein the historical data and real-time data include network status at different time periods and user feedback information.
[0193] In one embodiment, when the processor 502 implements the step of obtaining the WebRTC audio and video stream and the current network status, it specifically implements the following steps:
[0194] Create an RTCPeerConnection instance to obtain the WebRTC audio and video stream, and regularly call the getStats() method to obtain real-time statistical data of the WebRTC audio and video stream; extract key indicators from the real-time statistical data to obtain the current network status, where the key indicators include delay, packet loss rate, bandwidth, frame rate, resolution, and encoding format.
[0195] In one embodiment, when the processor 502 implements the step of performing comprehensive score calculation on the WebRTC audio and video stream to obtain a calculation result, the processor 502 specifically implements the following steps:
[0196] Obtain user feedback information on the WebRTC audio and video stream; calculate a comprehensive score using an indicator weighted average method according to the feedback information and the current network status to obtain a calculation result.
[0197] In one embodiment, when the processor 502 implements the step of obtaining the context model constructed by combining historical data and real-time data with the random forest algorithm for training, the adaptive model specifically implements the following steps:
[0198] Acquire historical data and real-time data; preprocess the historical data and the real-time data to obtain preprocessing results; perform feature selection on the preprocessing results to obtain selection results; divide the network state into different scenarios according to the selection results, and define corresponding labels and adjustment strategies to obtain a scenario model; use a random forest algorithm to train the scenario model to obtain an adaptive model.
[0199] In one embodiment, when the processor 502 implements the step of dynamically adjusting the buffer size of the audio and video data according to the adjustment strategy, the processor 502 specifically implements the following steps:
[0200] According to the adjustment strategy combined with the time series analysis method, the future network status change is predicted to dynamically adjust the buffer size of the audio and video data.
[0201] In one embodiment, when the processor 502 implements the step of adjusting the GOP size, QP value, and bit rate control according to the real-time monitored WebRTC audio and video streams and user preferences, the processor 502 specifically implements the following steps:
[0202] Determine the optimization priority level based on user preferences; combine the real-time monitored WebRTC audio and video streams to adjust the GOP size, QP value and bitrate control according to the optimization level.
[0203] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0204] It can be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment can be completed by instructing the relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiment of the above method.
[0205] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor executes the following steps:
[0206] Acquire WebRTC audio and video streams and current network status; perform comprehensive scoring calculation on the WebRTC audio and video streams to obtain a calculation result; when the calculation result is less than a set threshold, input the current network status into an adaptive model to obtain an adjustment strategy; dynamically adjust the buffer size of audio and video data according to the adjustment strategy; adjust the GOP size, QP value and bit rate control according to the real-time monitored WebRTC audio and video streams and user preferences.
[0207] The adaptive model is obtained by combining historical data and real-time data to build a situational model trained by a random forest algorithm, wherein the historical data and real-time data include network status at different time periods and user feedback information.
[0208] In one embodiment, when the processor executes the computer program to implement the step of obtaining the WebRTC audio and video stream and the current network status, the processor specifically implements the following steps:
[0209] Create an RTCPeerConnection instance to obtain the WebRTC audio and video stream, and regularly call the getStats() method to obtain real-time statistical data of the WebRTC audio and video stream; extract key indicators from the real-time statistical data to obtain the current network status, where the key indicators include delay, packet loss rate, bandwidth, frame rate, resolution, and encoding format.
[0210] In one embodiment, when the processor executes the computer program to implement the step of calculating a comprehensive score for the WebRTC audio and video stream to obtain a calculation result, the processor specifically implements the following steps:
[0211] Obtain user feedback information on the WebRTC audio and video stream; calculate a comprehensive score using an indicator weighted average method according to the feedback information and the current network status to obtain a calculation result.
[0212] In one embodiment, when the processor executes the computer program to implement the steps of constructing a situation model by combining historical data and real-time data with a random forest algorithm for training, the adaptive model specifically implements the following steps:
[0213] Acquire historical data and real-time data; preprocess the historical data and the real-time data to obtain preprocessing results; perform feature selection on the preprocessing results to obtain selection results; divide the network state into different scenarios according to the selection results, and define corresponding labels and adjustment strategies to obtain a scenario model; use a random forest algorithm to train the scenario model to obtain an adaptive model.
[0214] In one embodiment, when the processor executes the computer program to implement the step of dynamically adjusting the buffer size of the audio and video data according to the adjustment strategy, the processor specifically implements the following steps:
[0215] According to the adjustment strategy combined with the time series analysis method, the future network status change is predicted to dynamically adjust the buffer size of the audio and video data.
[0216] In one embodiment, when the processor executes the computer program to implement the step of adjusting the GOP size, QP value and bit rate control according to the real-time monitored WebRTC audio and video streams and user preferences, the processor specifically implements the following steps:
[0217] Determine the optimization priority level based on user preferences; combine the real-time monitored WebRTC audio and video streams to adjust the GOP size, QP value and bitrate control according to the optimization level.
[0218] The storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk, or any other computer-readable storage medium that can store program codes.
[0219] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0220] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0221] The steps in the method of the embodiment of the present invention can be adjusted in order, combined and deleted according to actual needs. The units in the device of the embodiment of the present invention can be combined, divided and deleted according to actual needs. In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0222] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, terminal, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.
[0223] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A method for dynamically adjusting WebRTC audio and video stream parameters, characterized in that: include: Get WebRTC audio and video streams and current network status; Performing a comprehensive score calculation on the WebRTC audio and video stream to obtain a calculation result; When the calculation result is less than a set threshold, the current network state is input into the adaptive model to obtain an adjustment strategy; Dynamically adjust the buffer size of audio and video data according to the adjustment strategy; Adjust the GOP size, QP value, and bitrate control based on real-time monitoring of WebRTC audio and video streams and user preferences.
2. The WebRTC audio and video stream parameter dynamic adjustment method according to claim 1, characterized in that: The obtaining of WebRTC audio and video streams and current network status includes: Create an RTCPeerConnection instance to obtain the WebRTC audio and video stream, and periodically call the getStats() method to obtain real-time statistics of the WebRTC audio and video stream; Key indicators are extracted from the real-time statistical data to obtain the current network status, wherein the key indicators include delay, packet loss rate, bandwidth, frame rate, resolution, and encoding format.
3. The method for dynamically adjusting WebRTC audio and video stream parameters according to claim 1, characterized in that: The performing a comprehensive score calculation on the WebRTC audio and video stream to obtain a calculation result includes: Obtain user feedback information on the WebRTC audio and video stream; A comprehensive score is calculated according to the feedback information and the current network status using an indicator weighted average method to obtain a calculation result.
4. The WebRTC audio and video stream parameter dynamic adjustment method according to claim 1, characterized in that: The adaptive model is obtained by combining historical data and real-time data to build a situational model trained by a random forest algorithm, wherein the historical data and real-time data include network status at different time periods and user feedback information.
5. The method for dynamically adjusting WebRTC audio and video stream parameters according to claim 4, characterized in that: The adaptive model is obtained by combining historical data and real-time data with a situational model trained by a random forest algorithm, including: Get historical data as well as real-time data; Preprocessing the historical data and the real-time data to obtain a preprocessing result; Performing feature selection on the preprocessing result to obtain a selection result; Dividing the network state into different scenarios according to the selection result, and defining corresponding labels and adjustment strategies to obtain a scenario model; The situational model is trained using a random forest algorithm to obtain an adaptive model.
6. The method for dynamically adjusting WebRTC audio and video stream parameters according to claim 1, characterized in that: The dynamically adjusting the buffer size of the audio and video data according to the adjustment strategy includes: According to the adjustment strategy combined with the time series analysis method, the future network status change is predicted to dynamically adjust the buffer size of the audio and video data.
7. The method for dynamically adjusting WebRTC audio and video stream parameters according to claim 1, characterized in that: The method of adjusting the GOP size, QP value and bitrate control according to the real-time monitored WebRTC audio and video stream and user preferences includes: Determine optimization priorities based on user preferences; Combined with real-time monitoring of WebRTC audio and video streams, the GOP size, QP value and bitrate control are adjusted according to the optimization level.
8. A WebRTC audio and video stream parameter dynamic adjustment device, characterized in that: include: The acquisition unit is used to obtain the WebRTC audio and video streams and the current network status; A calculation unit, configured to perform a comprehensive score calculation on the WebRTC audio and video stream to obtain a calculation result; An adjustment strategy generating unit, configured to input the current network state into an adaptive model to obtain an adjustment strategy when the calculation result is less than a set threshold; A first adjustment unit, configured to dynamically adjust the buffer size of the audio and video data according to the adjustment strategy; The second adjustment unit is used to adjust the GOP size, QP value and bit rate control according to the real-time monitored WebRTC audio and video streams and user preferences.
9. A computer device, characterized in that: The computer device comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Network state sensing and adaptive coding control method and device
CN120602051A
Data encryption transmission method and device for audio and video data, and electronic equipment
CN121309052A