Adaptive video stream optimization and transmission method and system

By dynamically adjusting the sharding time and pseudo-shatter filling mechanism, the problem of ignoring content complexity and fixed thresholds in video stream adaptive technology is solved, and efficient video streaming transmission and terminal optimization are achieved, ensuring picture quality and fluency.

CN120343328APending Publication Date: 2025-07-18CHINESE PEOPLES LIBERATION ARMY NAVAL ACAD

Patent Information

Application Number
CN202510480490.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing video streaming adaptive technology ignores the complexity of video content, resulting in blurring or block effects of scenes with high complexity, wasted bandwidth in low complexity scenes, and buffer management relies on fixed thresholds to easily cause lag or excessive image quality compression.

Method used

By dynamically adjusting the sharding time, pre-analyzing the scene complexity label, combining the device hardware data and physiological signals, dynamically selecting the decoding strategy, and triggering the pseudo-shatter filling mechanism when the buffer is below the threshold, optimizing video streaming transmission.

Benefits of technology

It realizes dynamic adjustments based on video content and network status, reduces picture lag, improves bandwidth utilization, ensures clarity of key information, balances picture quality and fluency, and adapts to the performance of different devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343328A_ABST
    Figure CN120343328A_ABST
Patent Text Reader

Abstract

The invention discloses an adaptive video stream optimization and transmission method and system, and the method comprises the steps: segmenting a video stream into fragments, and dynamically adjusting the fragment duration according to the proportion of a dynamic scene in the fragments; the method comprises the following steps: pre-analyzing a scene complexity label in head metadata of a fragment before decoding of a client, and dynamically selecting a decoding resolution and a frame rate of the fragment in combination with a current equipment screen refresh rate and environment illumination intensity sensor data; and in the transmission process, if it is detected that the buffer area occupancy rate of the client side is lower than a dynamic first threshold value based on the product of the network round-trip time and the fragment average code rate, a pseudo fragment filling mechanism is triggered. According to the invention, through scene intelligent perception, dynamic resource allocation and end-to-end cooperative control, the limitation of single-point optimization of a traditional scheme is broken through, and the problems caused by complexity difference of video contents and dependence of buffer management on a fixed threshold are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video stream adaptation, and in particular to an adaptive video stream optimization and transmission method and system. Background Art

[0002] With the popularization of mobile Internet and 5G technologies, video streaming media has become the core carrier in scenarios such as distance education, online conferencing, and entertainment consumption.

[0003] Existing video stream adaptation technologies, for example, a Chinese patent with the publication number CN108989838B discloses a DASH bitrate adaptation method based on video content complexity perception, belonging to the field of adaptive transmission of content distribution network videos. In this invention, the server end perceives the content complexity of the video to obtain the content complexity level of each segment in the video; the server end marks the content complexity level in the MPD file corresponding to the video; the client receives and utilizes the MPD file marked with the content complexity level to achieve DASH bitrate adaptation; this invention tries to avoid playback stagnation caused by buffer underflow when the network condition is bad, and avoid waste of network resources caused by buffer overflow when the network condition is good; it can distinguish the complexity of video segments to select the segment bitrate, and try to select a high-bitrate version for segments with high complexity. While improving the subjective viewing quality of segments with high complexity, it ensures that the viewing quality of all segments of the video is relatively high, thus providing a better quality of experience for users.

[0004] However, in the process of implementing the related technical solutions, it is found that at least the following technical problems exist:

[0005] 1. Only select the bitrate based on network bandwidth and buffer status, ignoring the complexity differences of the video content itself. For example: high-complexity scenarios (such as moving shots, pictures with rich texture details) require a higher bitrate to ensure clarity, but traditional methods may select a low bitrate due to network fluctuations, resulting in blurred pictures or block effects; low-complexity scenarios (such as static text, single background) cannot improve the subjective quality even if a high bitrate is allocated, but instead waste bandwidth.

[0006] 2. Buffer management depends on fixed thresholds, and buffer starvation (lag) or excessive compression of the picture quality (such as blurred text) is likely to occur when the network throughput suddenly changes or the device load suddenly increases. Summary of the Invention

[0007] To solve the above problems, an embodiment of the present invention provides an adaptive video stream optimization and transmission method, and the method includes:

[0008] When splitting the video stream into segments, dynamically adjust the segment duration according to the proportion of dynamic scenes within the segment;

[0009] Before client decoding, pre-analyze the scene complexity tags in the shard header metadata, and dynamically select the decoding resolution and frame rate of the shard in combination with the current device screen refresh rate and ambient light intensity sensor data;

[0010] During the transmission process, if it is detected that the occupancy rate of the client buffer is lower than the dynamic first threshold based on the product of the network round-trip time and the average bitrate of the shard, trigger the pseudo-shard filling mechanism.

[0011] Furthermore, the method for adjusting the shard duration includes:

[0012] When the dynamic scene ratio exceeds the adaptive threshold based on the historical motion characteristics of the shard, the shard duration is shortened to the first dynamic range; when the proportion of the static scene reaches the scene stability criterion, the shard duration is extended to the second dynamic range; the demarcation value between the first dynamic range and the second dynamic range is jointly determined by the current network bandwidth volatility and the device rendering ability;

[0013] The determination of the dynamic scene ratio includes:

[0014] Calculate the magnitude of the optical flow vector for each frame image in the shard, and count the proportion of pixel points whose magnitude exceeds the dynamic determination threshold corresponding to the device display resolution. If the average proportion of consecutive multiple frames exceeds the sliding window threshold based on the shard duration, mark it as a high dynamic scene shard; the sliding window threshold is inversely proportional to the device GPU processing ability.

[0015] Furthermore, according to the method described in claim 1, wherein the method of using the ambient light intensity sensor data includes:

[0016] When the ambient light intensity is lower than the dynamic adjustment threshold corresponding to the peak brightness of the device screen, reduce the video resolution according to the non-linear relationship between the light intensity and the screen brightness, and feedback to the server to adjust the HDR encoding strategy of subsequent shards, wherein the dynamic adjustment threshold is jointly calibrated according to the screen aging coefficient and the ambient color temperature data.

[0017] Furthermore, the pseudo-shard filling mechanism includes:

[0018] Send a low-bitrate pseudo-shard to maintain the buffer to the client, and preload the real shard in the background at the same time; when the buffer occupancy rate resumes to exceed η times of the dynamic first threshold, η>1 and the value of η is dynamically calculated according to the historical buffer recovery rate of the client, switch back to the real shard.

[0019] Furthermore, the pseudo-shard is a simplified version generated by using a semantic compression algorithm, retaining the high-precision encoding of the face area, and the background area is generated by contour interpolation based on edge detection, and its bitrate satisfies: bitrate = α×original shard bitrate + β×current bandwidth fluctuation variance, where α is negatively correlated with the device decoder complexity, and β is the network jitter rate compensation factor.

[0020] Furthermore, the adaptive video stream optimization and transmission method also includes physiological signal collaborative optimization of segmented transmission:

[0021] The time-frequency domain characteristics of the user's heart rate variability (HRV) are obtained through the client wearable device. When the standard deviation of the HRV data in the sliding time window exceeds the dynamic deviation threshold of the user's historical concentration state baseline value, the slice bit rate is reduced according to the exponential relationship between the deviation and the bit rate compression factor, and an interactive content trigger mark is injected into the video stream; the dynamic deviation threshold is adaptively updated according to the probability distribution of the user's recent physiological data.

[0022] Further, after the interactive content trigger mark is activated, if no tracking signal is detected within an adaptive time window based on the user's operating habits, the audio priority transmission mode is started;

[0023] In audio priority mode, the bitrate of the video segments is reallocated to improve the quality of the audio stream, and the allocation weight is determined by the convolution result of the sound source localization accuracy and the HRV deviation;

[0024] When the network bandwidth recovers to more than γ times the bandwidth required by the audio stream, the audio priority mode is automatically released.

[0025] Furthermore, the bit rate allocation strategy of the audio priority transmission mode works in coordination with the pseudo-slice filling mechanism, and the coordinated work includes:

[0026] When the audio priority transmission mode is activated, the pseudo-sliced semantic compression algorithm increases the voiceprint feature retention weight;

[0027] The video segment bitrate is reallocated to the resolution improvement of the audio stream, and its allocation ratio is determined by the convolution result of the user's HRV deviation and the sound source localization accuracy.

[0028] Adaptive video stream optimization and transmission system, the system includes:

[0029] A segmentation module, which dynamically adjusts the segment duration according to the dynamic scene ratio in the segment when segmenting the video stream into segments;

[0030] A pre-analysis module, which pre-analyzes the scene complexity tag in the metadata of the slice header before decoding on the client, and dynamically selects the decoding resolution and frame rate of the slice in combination with the current device screen refresh rate and ambient light intensity sensor data;

[0031] The transmission module triggers a pseudo-slice filling mechanism if it is detected during the transmission process that the client buffer occupancy rate is lower than a dynamic first threshold value based on the product of the network round-trip time and the average bit rate of the slices.

[0032] Technical effects and advantages of the adaptive video stream optimization and transmission method and system provided by the present invention:

[0033] Through scene intelligent perception, dynamic resource allocation, and end-to-end collaborative control, the present invention breaks through the single-point optimization limitation of traditional solutions and solves the problems brought by the complexity difference of the video content itself and the buffer management relying on fixed thresholds. By dynamically shortening the segment duration, the present invention significantly reduces the impact of sudden high bitrates on transmission stability and reduces frame freezes; by extending the segment duration, redundant requests are reduced, and the bandwidth utilization rate is significantly improved in low-dynamic content scenarios; by dynamically calibrating the segment strategy in combination with the terminal hardware capabilities, the rendering delay caused by excessive computational load on low-performance devices is avoided; by intelligently filling low-bitrate pseudo-segments, the playback continuity is maintained during network fluctuations, and the clarity of key visual information (such as faces, texts) is preferentially guaranteed; by dynamically adjusting the bitrate and rendering strategy according to the network state and device load, the requirements for image quality and smoothness are balanced; through full-link dynamic adaptation from segment segmentation, network transmission to terminal decoding, the collaborative optimization of smooth playback of high-dynamic content and low-power operation is achieved. Brief Description of the Drawings

[0034] Figure 1 It is a flowchart of the adaptive video stream optimization and transmission method in Embodiment 1;

[0035] Figure 2 It is a flowchart of the operation method after the interactive content trigger marker is activated in Embodiment 2;

[0036] Figure 3 It is a schematic connection diagram of the adaptive video stream optimization and transmission system in Embodiment 3. Detailed Embodiments

[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] Embodiment 1:

[0039] Please refer to Figure 1 As shown, the embodiment of the present invention provides an adaptive video stream optimization and transmission method, and the method includes:

[0040] When splitting the video stream into segments, dynamically adjust the segment duration according to the dynamic scene ratio within the segment;

[0041] Before client - side decoding, pre - analyze the scene complexity label in the shard header metadata, and dynamically select the decoding resolution and frame rate of the shard by combining the current device screen refresh rate and ambient light intensity sensor data;

[0042] During the transmission process, if it is detected that the occupancy rate of the client buffer is lower than the dynamic first threshold based on the product of the network round - trip time and the average bitrate of the shard, trigger the pseudo - shard filling mechanism.

[0043] The method for adjusting the shard duration includes:

[0044] When the dynamic scene ratio exceeds the adaptive threshold based on the historical motion characteristics of the shard, the shard duration is shortened to the first dynamic range; when the proportion of the static scene reaches the scene stability criterion, the shard duration is extended to the second dynamic range; the demarcation value between the first dynamic range and the second dynamic range is jointly determined by the current network bandwidth volatility and the device rendering ability.

[0045] In the video stream shard processing stage, the server executes a dynamic shard duration optimization strategy. Using the optical flow method to analyze the motion characteristics between consecutive frames, calculate the pixel - level optical flow vectors for all frames within each shard. For example, in a live basketball game scenario, when a player makes a quick breakthrough, the system detects that the optical flow amplitude of more than 60% of the pixel area is greater than the preset threshold, and at this time, the shard duration shortening mechanism is triggered. It should be noted that the adaptive threshold based on the historical motion characteristics of the shard is not a fixed value, but is dynamically adjusted by statistically calculating the median of the motion intensity of the previous N shards and combining the current network bandwidth jitter rate, where the value of N is automatically determined by the video content type classifier.

[0046] In terms of static scene processing, the implementation of the scene stability criterion includes multi - dimensional verification: first, use a convolutional neural network to identify the main stationary objects in the scene (such as buildings, trees, etc.), and then calculate the position offset variance of these objects within the shard time window. When this variance is lower than the camera jitter data collected by the device gyroscope, it is determined that the scene stability requirement is met; for example, in an outdoor surveillance scenario, when the camera is fixed and even if pedestrians pass by, as long as the background area meets the stability conditions, the system will still extend the shard duration to improve the encoding efficiency.

[0047] The specific method for determining the demarcation value between the first dynamic range and the second dynamic range includes:

[0048] Establish a weight relationship model between the network bandwidth volatility (calculating the coefficient of variation of the transmission rate of the last K shards through a sliding window) and the number of GPU shader cores of the device.

[0049] Exemplary: In the case of high bandwidth fluctuations (coefficient of variation > 0.3) in a 5G network and the terminal being equipped with a mobile GPU (such as Adreno 650), the demarcation value will automatically shift towards a shorter duration to ensure that the rendering delay does not exceed the human eye perception threshold.

[0050] The determination of the dynamic scene ratio includes:

[0051] Calculate the magnitude of the optical flow vector for each frame of the image within the slice, and count the proportion of pixel points whose magnitude exceeds the corresponding dynamic determination threshold of the device display resolution. If the average proportion of consecutive multiple frames exceeds the sliding window threshold based on the slice duration, it is marked as a high-dynamic scene slice; the sliding window threshold is inversely proportional to the processing capacity of the device GPU.

[0052] The implementation scheme of the corresponding dynamic determination threshold of the device display resolution includes:

[0053] Establish an optical flow magnitude conversion model based on the pixel density of the terminal screen.

[0054] For example, on a 4K resolution mobile device, since the physical size of the pixel points is smaller, the system will set an amplitude determination threshold 1.8 times higher than that of a 1080P device to avoid misjudging microscopic pixel jitter as effective motion. The dynamic adjustment of the sliding window threshold is achieved by monitoring the video memory bandwidth occupancy rate of the GPU. When the occupancy rate exceeds 75%, the determination conditions are automatically relaxed to reduce the computational load, specifically by adjusting the time coverage range of the sliding window.

[0055] Through the linkage of resolution adaptation and GPU load, the same algorithm automatically optimizes the calculation accuracy on different terminals; by combining object recognition and physical sensor data cross-verification, compared with the single image analysis method, it reduces; by coupling historical motion data with the real-time network state, it avoids the misjudgment problem caused by content mutation in the traditional scheme.

[0056] The usage method of the ambient light intensity sensor data includes:

[0057] When the ambient light intensity is lower than the corresponding dynamic adjustment threshold of the peak brightness of the device screen, reduce the video resolution according to the non-linear relationship between the light intensity and the screen brightness, and feedback to the server to adjust the HDR encoding strategy of the subsequent slices, where the dynamic adjustment threshold is jointly calibrated according to the screen aging coefficient and the ambient color temperature data.

[0058] The specific method for generating the dynamic adjustment threshold includes:

[0059] First, obtain the reference value of the screen peak brightness measured at the time of device factory shipment (such as 1000 nits for a certain OLED screen), then calculate the aging coefficient based on the cumulative value of the screen usage duration (such as a 3% brightness attenuation per 1000 hours of use), and at the same time combine the color coordinate data collected by the ambient color temperature sensor (such as the difference compensation between 6500K standard white light and 3000K warm light); for example, after the device has been used for two years, when it is detected that the actual peak brightness of the screen drops to 800 nits, dynamically calculate the adjustment threshold based on the current ambient color temperature (such as 4500K at dusk), so that it is reduced by 15%-20% compared to the new device state to adapt to the aging characteristics.

[0060] When the real-time ambient illuminance is lower than this dynamic threshold (such as only 50 lux of indoor light at night), the video resolution adjustment module starts the non-linear downscaling mechanism, which specifically includes:

[0061] Establish a hyperbolic tangent function relationship model between the illuminance (Lux) and the screen brightness (nits): Resolution scaling factor = 1 - α·tanh(β·(L max -L env ) / L max )), where L max is the current dynamic threshold, and α and β are shape parameters related to the device model; for example, on a tablet equipped with an AMOLED screen, when the ambient illuminance suddenly drops from 200 lux to 50 lux, the system automatically reduces the video resolution from 4K to 2K while maintaining high-definition rendering in the key area (such as within the face recognition frame).

[0062] After the resolution adjustment is completed, the client immediately sends an HDR policy update request carrying the environmental parameters to the server. The HDR encoding policy adjustment includes: dynamically selecting the mapping curve of the HDR metadata according to the current actual brightness capability of the screen (considering the aging coefficient) and the ambient color temperature data; for example, when the peak brightness of the screen is limited due to aging, use a compressed PQ curve to replace the standard ST2084 curve to avoid the loss of high-light details; the server performs pre-encoding optimization on the subsequent shards according to the feedback data, such as preferentially retaining the dark levels rather than the color saturation in low-light scenarios.

[0063] The pseudo-shard filling mechanism includes:

[0064] Send a low-bitrate pseudo-shard maintenance buffer to the client, and at the same time preload the real shards in the background; when the buffer occupancy rate resumes to exceed η times the dynamic first threshold (η > 1 and the η value is dynamically calculated according to the client's historical buffer recovery rate), switch back to the real shards.

[0065] Among them, the resolution selection strategy satisfies: the resolution is positively correlated with the square root of the screen refresh rate and negatively correlated with the logarithmic function of the illuminance;

[0066] The pseudo-fragment is a simplified version generated by a semantic compression algorithm, which retains the high-precision encoding of the face region and generates the background region by contour interpolation based on edge detection. Its bitrate satisfies: bitrate = α × original fragment bitrate + β × current bandwidth fluctuation variance, where α ∈ (0.1, 0.2) and is negatively correlated with the complexity of the device decoder, and β is the network jitter rate compensation factor.

[0067] For the processing of the face region, the YOLO-v5 neural network is used to detect the face region in the video frame in real time (when the detection box confidence is > 0.85), and the defined region (including a 20-pixel buffer zone around the facial feature points) is encoded using the B-frame inter prediction coding of the HEVC standard, and the quantization parameter (QP value) is set in the range of 18 - 22 to ensure that the facial texture detail retention rate > 95%; for example, in a video conferencing scenario, when multiple faces are detected, the system automatically allocates 60% of the bitrate for face region encoding.

[0068] Edge detection is performed on the non-face region, i.e., the background region (the Canny algorithm threshold is set to 50 - 150), the main contour structure of the scene is extracted through edge detection, and visual fidelity processing combining color block filling and gradient interpolation is used to generate a lightweight background version with an artistic style. In typical application scenarios, dynamic natural landscapes are converted into simple abstract pictures that retain the motion trend.

[0069] By adopting a differential processing strategy that differentiates the face core region from the background, the bitrate requirement is significantly reduced while ensuring the quality of the main picture; the background simplification algorithm combines the characteristics of human visual attention, making the picture quality change conform to the visual perception characteristics; the compression intensity adjustment mechanism sensitive to device capabilities extends the compatibility of different terminals; the bitrate compensation model responsive to network status enhances the transmission robustness in complex network environments.

[0070] Example 2:

[0071] This embodiment further improves the design on the basis of Embodiment 1. The difference is that in the actual operation of Embodiment 1, the problems of attention disconnection and network fluctuation sensitivity caused by fixed bitrate allocation are not considered, thus reducing the content transmission efficiency and user experience coherence in distance education and conferencing scenarios. Based on this, the adaptive video stream optimization and transmission method further includes the co-optimization of physiological signals for fragment transmission:

[0072] The time-frequency domain characteristics of the user's heart rate variability (HRV) are obtained through a client wearable device. When the standard deviation of the HRV data in the sliding time window exceeds the dynamic deviation threshold of the user's historical focused state baseline value, the fragment bitrate is reduced according to the exponential relationship between the deviation degree and the bitrate compression factor, and an interactive content trigger mark is injected into the video stream; the dynamic deviation threshold is adaptively updated according to the probability distribution of the user's recent physiological data.

[0073] Real-time obtain the photoplethysmogram (PPG) signal through wearable devices such as smartwatches, and adopt wavelet transform to extract the time-frequency domain features of heart rate variability (HRV). The time-frequency domain features include:

[0074] The ratio of low-frequency power (LF: 0.04 - 0.15 Hz) to high-frequency power (HF: 0.15 - 0.4 Hz);

[0075] The sliding window statistic of the standard deviation of RR intervals (SDNN).

[0076] The calculation method of the dynamic deviation threshold includes:

[0077] Construct a baseline library for the user's focused state:

[0078] Continuously collect the HRV features of the user during focused work (such as the efficient period from 10 - 11 am every day).

[0079] Establish the probability distribution of physiological features through the Gaussian mixture model;

[0080] When a change in the user's behavior pattern is detected (such as starting a fitness habit), adjust the baseline reference value using the exponentially weighted moving average method.

[0081] Bitrate dynamic adjustment mechanism:

[0082] Establish the mapping relationship between the HRV deviation (Δ) and the bitrate compression factor (K):

[0083] When Δ exceeds the threshold, trigger the non-linear compression of K = exp(-λΔ), where λ is the attenuation coefficient related to the content type;

[0084] Educational videos adopt progressive compression (retaining the clarity of the blackboard writing area);

[0085] Entertainment videos give priority to compressing high-frequency detail information.

[0086] Insert interactive markers at key frames, including:

[0087] The trigger area for the Q&A pop-up window (such as a floating multiple-choice question box that suddenly appears in the video);

[0088] The hot spot for gesture response (swipe a specific area of the screen to continue playing);

[0089] The voice interaction node (say the specified keyword to unlock the subsequent content).

[0090] Exemplary:

[0091] When the user's attention continuously deviates, the video automatically inserts an interactive checkpoint of "Please tap the screen to confirm understanding".

[0092] like Figure 2 As shown, after the interactive content trigger tag is activated, if no tracking signal (touch or eye movement) is detected within an adaptive time window based on the user's operating habits, the audio priority transmission mode is started;

[0093] According to the user's historical operating habits (such as the average reaction time of 2.3 seconds), the adaptive detection window is set to 3 seconds. If the touch click or gaze focus signal of the eye tracking device (such as the Tobii eye tracker) is not captured within this period, the client automatically switches to the audio priority transmission mode. For example, in the online language course scenario, if the user does not click the pop-up word test button in time, the system determines that he is in a non-active interaction state.

[0094] In audio priority mode, the bitrate of the video segments is reallocated to improve the quality of the audio stream, and the allocation weight is determined by the convolution result of the sound source localization accuracy and the HRV deviation;

[0095] After entering the audio priority mode, the bandwidth resources of the video segment are dynamically tilted towards the audio stream. The specific allocation weight is calculated by the sound source localization module and physiological data. The calculation method includes:

[0096] Assuming that a teaching video is transmitted, a microphone array is used to obtain the speaker's position (such as the teacher's voice positioning accuracy of ±15°), and its confidence score (0-1 standardized value) is convolved with the real-time HRV deviation; when the teacher is located in the center of the screen (confidence 0.9) and the HRV deviation is 1.8 times, the video bit rate is reduced from 3Mbps to 1Mbps, and the saved bandwidth is used to increase the audio sampling rate from 44.1kHz to 96kHz to achieve high-fidelity vocal enhancement.

[0097] When the network bandwidth recovers to more than γ times the bandwidth required by the audio stream (γ>1.2), the audio priority mode is automatically released.

[0098] The bandwidth recovery mechanism is triggered by real-time monitoring of network throughput. When it detects that the available bandwidth exceeds 1.3 times (γ=1.3) the current audio stream demand (such as 192kbps), the system automatically releases the audio priority mode. For example, after the network conditions improve, the video bit rate gradually recovers from 1Mbps to 3Mbps within a 2-second smooth transition period, while maintaining the audio enhancement state until the video quality is fully restored. This process avoids frequent mode switching through a dual-threshold hysteresis comparison algorithm to ensure a consistent user experience.

[0099] Exemplary:

[0100] Taking video conferencing as an example, when the user fails to respond to the interaction prompt in a timely manner due to fatigue during a remote meeting, the system activates audio optimization after 3 seconds: reduces the video resolution from 1080p to 720p, and at the same time enables the directional noise reduction function, increasing the speaker's voice signal-to-noise ratio by 12dB; when the network bandwidth rises from 1.5Mbps to 2.5Mbps (exceeding 1.3 times the audio stream demand), the video quality automatically recovers, avoiding communication interruptions caused by network fluctuations.

[0101] The bitrate allocation strategy of the audio priority transmission mode works in coordination with the pseudo-fragment filling mechanism:

[0102] When the audio priority transmission mode is activated, the semantic compression algorithm of the pseudo-fragments increases the weight of retaining voiceprint features;

[0103] The video fragment bitrate is reallocated to improve the resolution of the audio stream, and its allocation ratio is determined by the convolution result of the user's HRV deviation and the sound source localization accuracy.

[0104] When entering the audio priority transmission mode, the pseudo-fragment generation module starts semantic compression optimization. For the voice content in the video fragments, a deep learning-based voiceprint feature extraction algorithm is used to increase the weight of retaining the human voice features in the original audio stream from the base value (e.g., 0.6) to an increased value (e.g., 0.85). For example, in a remote medical consultation scenario, when the patient fails to respond to the doctor's questions in a timely manner, the system automatically strengthens the voiceprint features of the attending physician (such as fundamental frequency trajectory and formant distribution), increasing the speech intelligibility by 23%, and at the same time compressing the spectral energy of the background environmental noise to 40%.

[0105] The video bitrate reallocation process is implemented through a dynamic convolution kernel, which performs a time-domain convolution operation on the real-time calculated HRV deviation (such as the current SDNN value is 1.8 times the baseline value) and the sound source localization accuracy (such as the confidence of the speaker's position determined by the microphone array is 0.92). When the convolution result exceeds the threshold of 0.75, the bitrate of the video fragments is reallocated according to the gradient transfer rule, and the allocation mechanism includes:

[0106] Transfer 60% of the original video bitrate (such as 4Mbps) to enhance the audio stream quality; the video fragments enable adaptive pseudo-fragment filling and use the key frame interpolation algorithm to maintain the basic frame rate. For example, in an online teaching scenario, when the student's HRV deviation reaches 2.0 times and the teacher's localization accuracy is 0.95, the video bitrate drops from 4Mbps to 1.5Mbps, and the saved 2.5Mbps bandwidth is used to increase the audio bitrate from 128kbps to 320kbps and activate the 3D spatial audio rendering function.

[0107] The collaboration mechanism remains effective during the network recovery phase. When the bandwidth reaches the release condition (γ), the system first restores the quality of the audio stream enhanced by voiceprint, and then gradually increases the video bitrate. For example, during the process of the network bandwidth recovering from 2 Mbps to 3 Mbps, the audio sampling rate is first stabilized at 96 kHz / 24 bit, and then the video bitrate is linearly restored from 1.5 Mbps to 3 Mbps within 5 seconds, avoiding the perceptual discomfort caused by sudden changes in audio and video quality.

[0108] Exemplary:

[0109] In the scenario of a cross-border video conference, when the participants do not respond to the voting pop-up window in time due to network latency, the system activates the audio priority mode and, through the pseudo-fragmentation mechanism, increases the retention rate of the bitrate of the text area of the PPT shared screen to 80%, while compressing the background graphics to 40%.

[0110] Using the voiceprint feature enhancement algorithm, the MOS score of the speech of the main speaker is increased from 3.8 to 4.2.

[0111] According to the convolution result (HRV deviation 1.6 × positioning accuracy 0.87 = 1.39), 65% of the video bitrate is transferred to the audio stream. Under the bandwidth constraint of 384 kbps, the audio quality reaches the CD-level standard (44.1 kHz / 16 bit), while ensuring the readability of the key information of the presentation document.

[0112] Embodiment 3:

[0113] As Figure 3 shown, based on the same inventive concept as the adaptive video stream optimization and transmission method in the foregoing embodiment, the present application provides an adaptive video stream optimization and transmission system. The system in the embodiment of the present application and the method embodiment are based on the same inventive concept. Among them, the system includes:

[0114] A segmentation module that, when segmenting the video stream into fragments, dynamically adjusts the fragment duration according to the proportion of the dynamic scene within the fragment;

[0115] A pre-parse module that, before decoding by the client, pre-analyzes the scene complexity label in the fragment header metadata and dynamically selects the decoding resolution and frame rate of the fragment in combination with the current device screen refresh rate and ambient light intensity sensor data;

[0116] A transmission module that, during the transmission process, if it detects that the occupancy rate of the client buffer is lower than the dynamic first threshold based on the product of the network round-trip time and the average fragment bitrate, triggers the pseudo-fragment filling mechanism.

[0117] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

[0118] The above are only the preferred specific embodiments of the embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application, according to the technical solution and its concept of the present application, makes equivalent substitutions or changes, and should be covered by the protection scope of the present application.

Claims

1. An adaptive video stream optimization and transmission method, characterized in that Including: When splitting the video stream into segments, dynamically adjust the segment duration according to the proportion of the dynamic scene within the segment; Before decoding at the client side, pre-analyze the scene complexity label in the segment header metadata, and dynamically select the decoding resolution and frame rate of the segment in combination with the current device screen refresh rate and ambient light intensity sensor data; During the transmission process, if it is detected that the occupancy rate of the client buffer is lower than the dynamic first threshold based on the product of the network round-trip time and the average segment bit rate, trigger the pseudo-segment filling mechanism.

2. The adaptive video stream optimization and transmission method according to claim 1, wherein The method for adjusting the segment duration includes: When the proportion of the dynamic scene exceeds the adaptive threshold based on the historical motion characteristics of the segment, the segment duration is shortened to the first dynamic range; when the proportion of the static scene reaches the scene stability criterion, the segment duration is extended to the second dynamic range; the demarcation value between the first dynamic range and the second dynamic range is jointly determined by the current network bandwidth volatility and the device rendering ability; The determination of the proportion of the dynamic scene includes: Calculate the magnitude of the optical flow vector for each frame image within the segment, and count the proportion of pixel points whose magnitude exceeds the dynamic determination threshold corresponding to the device display resolution. If the average proportion of consecutive multiple frames exceeds the sliding window threshold based on the segment duration, mark it as a high-dynamic scene segment; the sliding window threshold is inversely proportional to the device GPU processing ability.

3. The adaptive video stream optimization and transmission method according to claim 1, wherein According to the method described in claim 1, wherein the method of using the ambient light intensity sensor data includes: When the ambient light intensity is lower than the dynamic adjustment threshold corresponding to the peak brightness of the device screen, reduce the video resolution according to the non-linear relationship between the light intensity and the screen brightness, and feedback to the server to adjust the HDR encoding strategy of subsequent segments, wherein the dynamic adjustment threshold is jointly calibrated according to the screen aging coefficient and the ambient color temperature data.

4. The adaptive video stream optimization and transmission method according to claim 1, characterized in that The pseudo-segment filling mechanism includes: Send a low-bitrate pseudo-segment to maintain the buffer to the client, and pre-load real segments in the background at the same time; when the buffer occupancy rate resumes to exceed η times of the dynamic first threshold, η>1 and the value of η is dynamically calculated according to the client's historical buffer recovery rate, switch back to real segments.

5. The adaptive video stream optimization and transmission method according to claim 4, characterized in that, The pseudo-segment is a simplified version generated by using a semantic compression algorithm, retaining high-precision encoding of the face area, and the background area is generated by contour interpolation based on edge detection. Its bit rate satisfies: bit rate = α×original segment bit rate + β×current bandwidth fluctuation variance, where α is negatively correlated with the device decoder complexity, and β is the network jitter rate compensation factor.

6. The adaptive video stream optimization and transmission method according to claim 1, wherein It also includes the collaborative optimization of physiological signals for segment transmission: Obtain the time-frequency domain characteristics of the user's heart rate variability (HRV) through the client's wearable device. When the standard deviation of the HRV data within the sliding time window exceeds the dynamic deviation threshold of the user's historical concentration state baseline value, reduce the segment bit rate according to the exponential relationship between the deviation degree and the bit rate compression factor, and inject an interactive content trigger mark into the video stream; The dynamic deviation threshold is adaptively updated according to the probability distribution of the user's recent physiological data.

7. The adaptive video stream optimization and transmission method according to claim 6, wherein After the interactive content trigger mark is activated, if no tracking signal is detected within the adaptive time window based on the user's operation habits, start the audio priority transmission mode; In the audio - priority mode, the bitrate of video segments is re - allocated to improve the quality of the audio stream, and the allocation weight is determined by the convolution result of the sound - source localization accuracy and the HRV deviation; When the network bandwidth recovers to more than γ times the required bandwidth of the audio stream, the audio - priority mode is automatically lifted.

8. The adaptive video stream optimization and transmission method according to claim 7, wherein The bitrate allocation strategy of the audio - priority transmission mode works in coordination with the pseudo - segment filling mechanism. The coordination includes: When the audio - priority transmission mode is activated, the semantic compression algorithm of pseudo - segments increases the weight of retaining voiceprint features; The bitrate of video segments is re - allocated to improve the resolution of the audio stream, and the allocation ratio is determined by the convolution result of the user's HRV deviation and the sound - source localization accuracy.

9. Adaptive video stream optimization and transmission system, characterized in that The system includes: A segmentation module that, when segmenting the video stream into segments, dynamically adjusts the segment duration according to the proportion of dynamic scenes within the segment; A pre - parsing module that, before decoding on the client side, pre - analyzes the scene - complexity label in the segment - header metadata and dynamically selects the decoding resolution and frame rate of the segment in combination with the current device screen refresh rate and ambient light - intensity sensor data; A transmission module that, during the transmission process, if it detects that the occupancy rate of the client buffer is lower than a dynamic first threshold based on the product of the network round - trip time and the average bitrate of the segment, triggers the pseudo - segment filling mechanism.

Citation Information

Patent Citations

  • A DASH bitrate adaptive method based on video content complexity awareness

    CN108989838B

Cited By

  • Ultrahigh-definition video stream multi-mode low-delay division method based on dynamic resource scheduling

    CN120751136A