5g-based visual call information processing method and system
By constructing a network transmission adaptation model and content-aware encoding conversion, the transmission efficiency and quality issues in visual call information processing under 5G environment were solved, achieving efficient and stable transmission of video data and decoding timing alignment, thus improving call fluency.
Patent Information
- Application Number
- CN202511492089.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-10-20
AI Technical Summary
Traditional visual call information processing methods struggle to balance transmission efficiency and content quality in a 5G environment. Delayed network status response leads to data loss or transmission delays, and the lack of timing coordination mechanisms in the decoding stage affects call fluency.
By acquiring real-time video frame sequences and 5G network environment perception data, a network transmission adaptation model is constructed, video encoding control instructions are generated, content-aware encoding conversion is performed, dynamic calibration is carried out during transmission, decoding timing alignment is processed, and a synchronized visual call output sequence is generated.
It achieves efficient and stable transmission of video data in a 5G network environment, ensuring the real-time performance and quality of video calls, and improving the dynamic adaptation capability between network status and video encoding.
Smart Images

Figure CN120980184B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and more specifically, to a method and system for processing visual call information based on 5G. Background Technology
[0002] With the development of 5G communication technology, video calling, as an important application in real-time interactive scenarios, places high demands on the stability of video data transmission, content quality, and time consistency. Traditional video call information processing methods suffer from insufficient adaptability between encoding strategies and network conditions, making it difficult to balance transmission efficiency and content quality during network fluctuations. Furthermore, the response of transmission strategy adjustments to link characteristics is lagging, easily leading to data loss or transmission delays when bandwidth suddenly drops or latency suddenly increases. The decoding stage lacks a timing coordination mechanism with the encoding and transmission stages, potentially affecting call fluency due to accumulated timing deviations. The overall processing flow fails to achieve dynamic linkage between network conditions, content characteristics, and transmission / decoding, making it difficult to meet the comprehensive requirements of real-time performance, stability, and quality for video calls in the 5G environment. Summary of the Invention
[0003] In view of this, the present invention provides a 5G-based method and system for processing visual call information.
[0004] According to one aspect of the present invention, a 5G-based visual call information processing method is provided. The method includes: continuously acquiring real-time video frame sequences and 5G network environment perception data in a visual call scenario, wherein the real-time video frame sequences form video stream units in timestamp order, and the 5G network environment perception data reflects the dynamic change characteristics of the current transmission link; performing multi-dimensional state mapping processing on the 5G network environment perception data to construct a network transmission adaptation model, wherein the network transmission adaptation model is used to characterize the dynamic correlation between network state and video encoding requirements; generating video encoding control instructions based on the network transmission adaptation model, performing content-aware encoding conversion on the real-time video frame sequences, and outputting an optimized encoding stream; during the transmission of the optimized encoding stream, acquiring link state fluctuation information through a 5G network feedback channel, dynamically calibrating the transmission strategy of the optimized encoding stream based on the link state fluctuation information, and obtaining a calibrated transmission stream; performing decoding timing alignment processing on the calibrated transmission stream to generate a visual call output sequence synchronized with the original video stream units, and pushing it to a receiving end presentation device.
[0005] According to another aspect of the present invention, a computer system is provided, comprising: a processor; and a memory, wherein the memory stores computer-readable code, which, when executed by the processor, causes the processor to perform the 5G-based video call information processing method as described above. Attached Figure Description
[0006] Figure 1 This is a schematic diagram of the application scenario provided by the present invention;
[0007] Figure 2 This is a flowchart illustrating a 5G-based visual call information processing method provided by the present invention.
[0008] Figure 3 This is a schematic diagram of the structure of a computer system provided in an embodiment of the present invention. Detailed Implementation
[0009] To facilitate a clearer understanding of this invention, we will first introduce the application scenarios in which this invention is implemented, such as... Figure 1 As shown, the system includes a computer system 10 and a terminal cluster. The terminal cluster may include one or more terminals; the number of terminals is not limited here. Figure 1 As shown, the terminal cluster may specifically include terminal 1, terminal 2, ..., terminal n; it can be understood that terminal 1, terminal 2, terminal 3, ..., terminal n can all be connected to the computer system 10 via a network so that each terminal can interact with the computer system 10 via the network connection.
[0010] Understandably, computer system 10 can refer to a device executing the method of this invention, and can be a server. A server can be a single physical server, a server cluster or distributed system consisting of at least two physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can specifically refer to a presentation device performing 5G video calls, such as an in-vehicle terminal, smartphone, tablet, laptop, desktop computer, etc., but is not limited to these. Each terminal and the server communicate directly via 5G.
[0011] Further, please see Figure 2 This is a flowchart illustrating a 5G-based visual call information processing method provided in an embodiment of the present invention. Figure 2 As shown, this method can be derived from... Figure 1 The computer system in the process executes the 5G-based visual call information processing method, which may include the following steps:
[0012] Step S100: Continuously acquire real-time video frame sequences and 5G network environment perception data in the visual call scenario. The real-time video frame sequences form video stream units in timestamp order, and the 5G network environment perception data reflects the dynamic change characteristics of the current transmission link.
[0013] A real-time video frame sequence is a series of video frames captured sequentially in chronological order during a video call. These frames contain image information from both parties in the call. A timestamp is a time identifier for each video frame, used to determine the order and playback time of the video frames. A video stream unit is composed of these real-time video frame sequences arranged in timestamp order, forming a continuous unit of video data. 5G network environmental awareness data is obtained through real-time monitoring of the current transmission link of the 5G network using designated sensors and monitoring equipment. It reflects the dynamic changes of the network at different times, such as fluctuations in network bandwidth, changes in signal strength, and packet loss rate.
[0014] Step S200: Perform multi-dimensional state mapping processing on the 5G network environment perception data to construct a network transmission adaptation model. The network transmission adaptation model is used to characterize the dynamic correlation between network state and video encoding requirements.
[0015] Multi-dimensional state mapping processing analyzes and processes 5G network environment perception data from multiple different angles and dimensions, mapping it to different network states. The network transmission adaptation model is a mathematical model describing the relationship between network states and video encoding requirements. It can dynamically adjust video encoding parameters and strategies based on the current network state to ensure efficient and stable video transmission in different network environments.
[0016] In some implementations, step S200 may include the following steps S210 to S250:
[0017] Step S210: Collect wireless resource occupancy and signal quality change trends from 5G network environment perception data, establish a network status time series database, and store network characteristic variables according to the sampling period in the network status time series database.
[0018] Wireless resource occupancy refers to the degree and status of the use of various wireless resources (such as spectrum resources and channel resources) in a 5G network. Signal quality change trend refers to the change in the quality of 5G network signals (such as signal strength and signal-to-noise ratio) over time. The network state time series database is a database used to store network state data. It stores collected network characteristic variables (such as wireless resource occupancy rate and signal strength) according to a certain sampling period (such as once per second) for subsequent analysis and processing.
[0019] In practice, the collection of wireless resource occupancy information can be achieved through the resource management module of the 5G base station, which can monitor and statistically analyze the usage of various wireless resources in real time, such as the occupancy time and bandwidth of each channel. For collecting signal quality change trends, signal strength sensors and signal-to-noise ratio (SNR) monitoring devices can be used to monitor the strength and SNR of the 5G network signal in real time and record their changes over time. Then, this collected data is stored in a network status time-series database according to a sampling period (e.g., once per second).
[0020] Step S220: Perform spatiotemporal correlation analysis on the data in the network state time series database to identify the periodic patterns and sudden disturbance characteristics of network state fluctuations and generate a network state feature map.
[0021] Spatiotemporal correlation analysis analyzes data from a network state time-series database from both temporal and spatial dimensions to study the interrelationships and patterns of change among the data. The periodicity of network state fluctuations refers to the cyclical changes in network state over a certain time frame; for example, network bandwidth exhibits regular fluctuations within specific time periods each day. Sudden disturbances are abrupt and abnormal changes in network state at certain moments, such as a sudden drop in network bandwidth due to sudden network congestion or interference. A network state feature map is a visual representation of network state characteristics, clearly presenting the periodicity and sudden disturbances of network state through graphical and data-driven methods.
[0022] In some implementations, step S220 may include the following steps S221 to S225:
[0023] Step S221: Perform time series decomposition on the data in the network state time series database to separate the trend term, periodic term and random disturbance term. The trend term reflects the long-term change direction of the network state, the periodic term reflects the periodic fluctuation characteristics of the network state, and the random disturbance term represents the state change caused by sudden factors.
[0024] Time series decomposition is a method of breaking down time series data into different components. This method can decompose data in a network state time series database into trend components, periodic components, and random disturbance components. The trend component represents the overall trend of network state changes over a longer period, such as network bandwidth gradually increasing over time. The periodic component represents the periodic fluctuations in network state within a certain time range, such as regular fluctuations in network bandwidth during specific time periods each day. The random disturbance component represents sudden changes in network state caused by unexpected factors (such as network failures, interference, etc.); these changes are random and unpredictable.
[0025] In practice, classic time series decomposition methods, such as additive or multiplicative models, can be used. For separating the trend term, the moving average method can be used to smooth data fluctuations and obtain the long-term trend of the network state by calculating the average value of data within a certain time window. For separating the periodic term, methods such as Fourier transform can be used to convert the time-domain signal into a frequency-domain signal and identify the main periodic components. For separating the random disturbance term, it can be obtained by subtracting the trend and periodic terms from the original data.
[0026] Step S222: Perform spectral analysis on the periodic term to determine the main periodic components and corresponding amplitude characteristics of the network state fluctuations. Convert the time-domain signal into a frequency-domain signal through Fourier transform and identify the peak frequency in the spectrum as the main periodic component.
[0027] Spectrum analysis is an analytical method that converts time-domain signals into frequency-domain signals. This method allows for the separation of different frequency components within a periodic term and the determination of the amplitude and phase of each component. The dominant periodic component is the one that has the greatest impact on network state fluctuations within the periodic term; its corresponding frequency is the peak frequency in the spectrum. The amplitude characteristic represents the magnitude of the dominant periodic component, reflecting the degree of its influence on network state fluctuations.
[0028] In practice, the separated periodic data is first used as input and converted into a frequency domain signal using a Fourier transform algorithm. The Fourier transform algorithm calculates the amplitude and phase of each frequency component, forming a spectrum. Then, the frequency with the largest amplitude is found in the spectrum; this frequency is the frequency of the main periodic component. Simultaneously, the amplitude corresponding to this frequency is recorded as the amplitude characteristic of the main periodic component.
[0029] Step S223: Perform statistical distribution analysis on the random disturbance term, calculate the probability density function and cumulative distribution function of the disturbance amplitude, determine the threshold condition for the occurrence of the disturbance, and determine the sudden disturbance event when the disturbance amplitude exceeds the threshold.
[0030] Statistical distribution analysis is a method for studying and analyzing the probability distribution of data, allowing us to understand the distribution characteristics of random disturbance terms. The probability density function describes the probability distribution of a random variable around a certain value, reflecting the likelihood of the variable taking that value. The cumulative distribution function describes the probability that a random variable is less than or equal to a certain value. The disturbance amplitude is the magnitude of change in the random disturbance term, i.e., the degree of deviation of the random disturbance term from its normal state. The threshold condition is a critical value determined based on statistical analysis results; when the disturbance amplitude exceeds this threshold, it is considered a sudden disturbance event.
[0031] In practice, the separated random disturbance term data is first statistically analyzed to calculate its mean, variance, and other statistical parameters. Then, based on these statistical parameters, the probability density function and cumulative distribution function of the random disturbance term are fitted. Common probability distribution models, such as the normal distribution and Poisson distribution, can be used to fit the distribution of the random disturbance term. For example, by analyzing a large amount of historical random disturbance term data, it is found that it conforms to a normal distribution; therefore, the probability density function and cumulative distribution function of the normal distribution can be used to describe the distribution characteristics of the random disturbance term. Next, a suitable threshold condition is determined based on the probability density function and cumulative distribution function.
[0032] Step S224: Combine the trend term, main periodic components and sudden disturbance events to construct the spatiotemporal dimensions of the network state feature map. The time dimension is marked with the network state change curve in units of sampling period, and the spatial dimension is plotted with network resource occupancy and signal quality as coordinate axes to draw a scatter plot of state distribution.
[0033] The spatiotemporal dimension of the network state feature map displays the characteristics of the network state through two dimensions: time and space. The time dimension, using the sampling period as a unit, plots the changes in the network state as a curve, reflecting the trend of network state changes over time. The spatial dimension, using network resource utilization and signal quality as coordinate axes, plots the network state at different times as scattered points on a two-dimensional plane, reflecting the spatial distribution of the network state.
[0034] Step S225: Mark the periodic regular regions and sudden disturbance regions in the network state feature map. The periodic regular regions are defined by the confidence interval of the periodic components, and the sudden disturbance regions are marked by the occurrence time and impact range of the disturbance event, so as to obtain the complete network state feature map.
[0035] Periodicity regions are areas within which the network state falls within a certain confidence interval during periodic fluctuations. A confidence interval is a statistical concept used to represent the range of values for a parameter with a certain probability. In the network state feature map, periodicity regions are defined by the confidence intervals of the periodic components, reflecting the normal range of changes in the network state during periodic fluctuations. Sudden disturbance regions are areas where the network state is affected by sudden disturbance events. They are marked by the timing and extent of the disturbance event, clearly demonstrating the impact of sudden disturbance events on the network state.
[0036] In practice, for marking periodic regions, the confidence interval of the periodic component is first determined based on the statistical analysis results of the periodic term. In the network state feature map, regions falling within this confidence interval are marked as periodic regions. For marking sudden disturbance regions, when a sudden disturbance event is detected, the time of occurrence and the affected area are recorded. The affected area can be determined by the magnitude and duration of the network state change.
[0037] Step S230: Construct a state transition probability matrix based on the network state feature map. The state transition probability matrix is used to describe the probability of transitions and the duration distribution between different network states.
[0038] The state transition probability matrix is a probability matrix used to describe the transition of network states between different states. Each row and column corresponds to a network state, and each element in the matrix represents the probability of transitioning from one network state to another. The duration distribution is the distribution of the network's duration in each state, which can be obtained through statistical analysis of historical data.
[0039] In some implementations, step S230 may include the following steps S231 to S235:
[0040] Step S231: Divide the network state in the network state feature map into multiple discrete state levels. Each state level corresponds to a set resource occupancy rate range and signal quality interval. Use a clustering algorithm to classify the continuous network state data and determine the division boundary of the state level.
[0041] Discrete state hierarchy divides continuous network state data into several discrete state categories, each corresponding to a defined range of resource occupancy and signal quality. Clustering algorithms are unsupervised learning algorithms used to group similar data points together into different categories. When constructing the state transition probability matrix, the continuous network state data in the network state feature map needs to be discretized to facilitate subsequent probability calculations. In practice, the K-Means clustering algorithm can be used to classify the network state data in the network state feature map.
[0042] Step S232: Count the number of transitions between different state levels, calculate the probability value of transitioning from one state level to another, the probability value is the ratio of the number of transitions to the total number of transitions, and obtain the initial state transition probability matrix.
[0043] The number of transitions is the number of times the network state transitions from one state level to another within a certain period. The total number of transitions is the sum of the number of transitions between all state levels. The initial state transition probability matrix is a probability matrix calculated based on the statistically obtained number of transitions and the total number of transitions, reflecting the probability of the network state transitioning between different state levels. In actual statistical processing, it is necessary to analyze historical network state data, record the network state level at each moment, and count the number of transitions between different state levels. After counting all the transitions, the probability value of transitioning from one state level to another is calculated. The transition probability values between all state levels are then filled into the matrix to obtain the initial state transition probability matrix.
[0044] Step S233: Analyze the duration distribution of each state level, calculate the average duration and variance of the state level through survival analysis, and establish a duration distribution model. The duration distribution model is used to describe the stability characteristics of the state level in the time dimension.
[0045] Survival analysis is a statistical method used to study the timing of events. It can be used to analyze the duration distribution of network states at each state level. The average duration is the average length of time a network state lasts at a given state level, while the variance reflects the dispersion of the duration. A duration distribution model is a model used to describe the stability characteristics of state levels over time, and it can be obtained through statistical analysis of historical data.
[0046] In practical analysis, for each state level, the times when a network state enters and leaves that state level are recorded, and the duration of each state level is calculated. Then, survival analysis methods (such as the Kaplan-Meier estimation method) are used to analyze this duration data, calculating the mean duration and variance. Based on these statistical results, a duration distribution model is established, for example, using an exponential or gamma distribution to fit the duration data, resulting in a duration distribution model. This model can be used to predict the duration of a network state at a given state level and to assess the stability of that state level.
[0047] Step S234: Combining the initial state transition probability matrix and the duration distribution model, construct a state transition probability matrix containing two parameters: transition probability and duration. Each element in the matrix contains the transition probability from the current state to the target state and the average duration in the current state.
[0048] Based on the initial state transition probability matrix, and combined with the duration distribution model, a state transition probability matrix with two parameters, transition probability and duration, is constructed. Each element in this matrix contains not only the transition probability from the current state to the target state, but also the average duration in the current state.
[0049] Step S235: Normalize the state transition probability matrix so that the sum of the probabilities of each row is 1. Verify the convergence of the matrix using a Markov chain model. The construction is complete when the matrix satisfies the stationary distribution condition.
[0050] Normalization involves adjusting the elements of the state transition probability matrix so that the sum of the probabilities of each row is 1. This is because during a state transition, the sum of the probabilities of transitioning from one state to all other possible states should be 1. Markov chain models are probability-based stochastic process models used to describe the transitions of a system between different states. Convergence is whether the matrix can reach a stable state after multiple iterations. The stationary distribution condition states that after the matrix reaches a stable state, the probability distribution of each state no longer changes.
[0051] Step S240: Associate the state transition probability matrix with the preset video coding requirement index to generate the basic framework of the network transmission adaptation model. The basic framework includes a network state input layer, a feature mapping layer, and a coding requirement output layer.
[0052] Pre-defined video coding requirements are a series of metrics pre-set based on the characteristics and transmission requirements of the video, such as video resolution, frame rate, and bit rate. Association modeling associates the state transition probability matrix with these video coding requirements, establishing a mathematical relationship between them. The basic framework of the network transmission adaptation model is a three-layer structure model comprising a network state input layer, a feature mapping layer, and a coding requirement output layer. The network state input layer receives current network state information, the feature mapping layer processes and transforms the input network state information, and the coding requirement output layer outputs the video coding requirements determined based on the network state.
[0053] Step S250: Dynamically adjust the basic framework using historical transmission data to optimize the mapping relationship between network status and encoding requirements, so that the network transmission adaptation model can reflect the video encoding adaptation requirements under the current network environment in real time.
[0054] Historical transmission data records network status information, corresponding video encoding requirements, and transmission performance (such as video quality and packet loss rate) during past network transmissions. Dynamic calibration adjusts and optimizes the basic framework of the network transmission adaptation model based on historical transmission data, enabling the model to better adapt to different network environments and video encoding requirements. The mapping relationship is the correspondence between network status and video encoding requirements; dynamic calibration can make this mapping relationship more accurate and reasonable.
[0055] In the actual calibration process, a large amount of historical transmission data is first collected, including network status information (such as bandwidth, signal strength, etc.), video encoding requirement indicators (such as resolution, frame rate, bit rate, etc.), and transmission performance data (such as video quality score, packet loss rate, etc.). Then, machine learning algorithms (such as linear regression, neural networks, etc.) are used to analyze and model the historical transmission data. The parameters of the model are obtained by training on the historical data. Next, these parameters are used to adjust the basic framework of the network transmission adaptation model, optimizing the mapping relationship between network status and encoding requirements.
[0056] Step S300: Generate video encoding control instructions based on the network transmission adaptation model, perform content-aware encoding conversion on the real-time video frame sequence, and output the encoded optimized stream.
[0057] Video encoding control instructions are generated based on video encoding requirement indicators output by the network transmission adaptation model. These instructions control the video encoder's encoding operations and include information such as encoding mode, quantization parameters, and bitrate control. Content-aware encoding conversion employs different encoding strategies based on the content characteristics of video frames (such as regions of interest and non-regions of interest) to improve encoding efficiency and video quality. Encoded optimized streams are the video streams obtained after content-aware encoding conversion. They minimize data volume and improve transmission efficiency while maintaining video quality.
[0058] In some implementations, step S300 may include the following steps S310 to S360:
[0059] Step S310: Analyze the encoding requirement index output by the network transmission adaptation model to determine the video encoding priority parameter under the current network condition. The encoding priority parameter is used to guide the encoding resource allocation strategy for different video content.
[0060] Encoding requirement metrics are a series of video encoding parameters output by the network transmission adaptation model based on the current network state, such as video resolution, frame rate, and bitrate. Video encoding priority parameters are determined based on these encoding requirement metrics and are used to distinguish the importance of different video content, thereby guiding the allocation of encoding resources. Encoding resource allocation strategies, based on the video encoding priority parameters, determine how to allocate encoding resources (such as bitrate and quantization step size) to different video content.
[0061] In the actual analysis process, the encoding requirement indicators output by the network transmission adaptation model are first analyzed and interpreted. Then, based on these encoding requirement indicators, video encoding priority parameters are determined. A mapping relationship can be established to assign different combinations of encoding requirement indicators to different video encoding priority parameters. Based on the video encoding priority parameters, an encoding resource allocation strategy is formulated. For higher-priority video content, more encoding resources are allocated, such as a higher bitrate and a smaller quantization step size; for lower-priority video content, fewer encoding resources are allocated, such as a lower bitrate and a larger quantization step size.
[0062] Step S320: Perform content saliency analysis on the real-time video frame sequence to identify regions of interest and non-regions of interest in the video frames. Regions of interest are determined by a visual attention model, and non-regions of interest are areas in the video frames whose saliency values are below a threshold.
[0063] Content saliency analysis analyzes the content within video frames to determine the saliency of different regions. Saliency reflects the visual attractiveness of a region and is typically related to its color, texture, motion, and other features. Regions of interest (ROIs) are important areas in a video frame that a user is likely to focus on; these are determined using visual attention models. Visual attention models simulate human visual attention mechanisms and can predict regions that humans are likely to focus on based on the characteristics of a video frame. Regions of non-interest (ROIs) are areas in a video frame with saliency values below a certain threshold; these areas have a relatively small impact on the user's visual experience.
[0064] In some implementations, step S320 may include the following steps S321 to S326:
[0065] Step S321: Extract color features from each video frame in the real-time video frame sequence, convert the video frame from the RGB color space to the LAB color space, separate the luminance component and chrominance component, and calculate the standard deviation of the chrominance component as the color contrast feature.
[0066] Color feature extraction involves extracting color-related feature information from video frames. The RGB color space is a common color representation method, using three channels: red (R), green (G), and blue (B). The LAB color space is more consistent with human visual perception, dividing color information into two parts: luminance (L) and chrominance (a, b). The luminance component reflects the brightness of a color, while the chrominance component reflects its hue and saturation. Color contrast features are obtained by calculating the standard deviation of the chrominance components, reflecting the degree of color variation within a video frame.
[0067] Step S322: Extract the texture features of the video frame, and use the gray-level co-occurrence matrix to calculate the energy, entropy, contrast and correlation parameters of the texture. The texture features are used to characterize the distribution of detailed areas in the video frame.
[0068] Texture features are the characteristic information of texture in a video frame, reflecting the distribution of detailed regions within the frame. The gray-level co-occurrence matrix (GLCM) is a statistical method used to describe image texture. It extracts texture feature parameters by calculating the co-occurrence relationships between different gray levels in an image. Energy, entropy, contrast, and correlation are four important texture feature parameters calculated from the GLCM. Energy reflects the uniformity of the texture, entropy reflects its randomness, contrast reflects its sharpness, and correlation reflects its local similarity.
[0069] Step S323: Analyze the motion characteristics of the video frames, calculate the pixel displacement vector between adjacent video frames using an optical flow estimation algorithm, and identify regions whose motion vector magnitude exceeds a preset motion threshold as potential regions of interest.
[0070] Motion features are the motion information of objects in a video frame, reflecting their dynamic changes. Optical flow estimation algorithms are methods used to calculate pixel displacement vectors between adjacent video frames. By analyzing the brightness changes of pixels in adjacent video frames, they estimate pixel motion. The pixel displacement vector represents the direction and magnitude of a pixel's displacement in adjacent video frames. A preset motion threshold is a pre-defined value used to determine whether pixel motion is significant. Potential regions of interest (ROIs) are areas in a video frame where the magnitude of the motion vector exceeds the preset motion threshold. These areas are typically regions where object motion is more pronounced and may be of significant interest to the user. In practical analysis, optical flow estimation algorithms (such as the Lucas-Kanade algorithm) can be used to calculate pixel displacement vectors between adjacent video frames.
[0071] Step S324: Input color contrast features, texture features and motion features into the visual attention model, and fuse them by weighted summation. The weight coefficient of motion features in potential regions of interest is set higher than that of motion features outside the region to generate a saliency map of the video frame. The value of each pixel in the saliency map represents the visual saliency of that pixel.
[0072] Visual attention models simulate the human visual attention mechanism, predicting regions that humans might focus on based on input feature information. Weighted summation is a method of fusing multiple features by assigning a weight coefficient to each feature and summing them to obtain a comprehensive feature value. A saliency map is a matrix the same size as a video frame, where the value of each pixel represents its visual saliency, ranging from [0,1], with larger values indicating more saliency.
[0073] In the actual fusion process, color contrast features, texture features, and motion features are first input into the visual attention model. The visual attention model can be a machine learning-based model, such as a convolutional neural network (CNN). During the weighted summation, a weight coefficient is assigned to each feature. Simultaneously, the weight coefficients for motion features within potential regions of interest are set higher than those outside these regions. Then, the color contrast features, texture features, and motion features of each pixel are weighted and summed to obtain the saliency value of that pixel. The saliency values of all pixels are combined to generate a saliency map of the video frame.
[0074] Step S325: Binarize the saliency map, mark the pixel regions with saliency values higher than the threshold as regions of interest, and mark the regions with saliency values lower than the threshold as regions of non-interest. The threshold is determined by the maximum inter-class variance method.
[0075] Binarization is the process of converting pixel values in a saliency map into binary values (0 or 1). This process can divide video frames into regions of interest (ROI) and non-ROI. The maximum inter-class variance (MOV) method is used to determine the binarization threshold. By maximizing the inter-class variance, it finds an optimal threshold that maximizes the difference between the two classes (foreground and background) when dividing the image. In practical processing, the MOV method can be used to determine the binarization threshold.
[0076] Step S326: Perform morphological processing on the marked region of interest to remove noise regions with areas smaller than a preset value, and optimize the region boundary through dilation and erosion operations to ensure the integrity and continuity of the region of interest.
[0077] Morphological processing is a shape-based image processing method that alters the shape and structure of an image by performing operations such as dilation, erosion, opening, and closing. Noise regions are small areas within the marked region of interest (ROI) that may be caused by noise or misclassification. The preset value is a pre-defined area value used to determine whether a region is noise. Dilation expands the foreground region outwards, while erosion contracts it inwards. Through dilation and erosion, the boundaries of the ROI can be optimized, making the region more complete and continuous.
[0078] Step S330: Based on the coding priority parameters and content saliency analysis results, generate a differentiated coding scheme. A high-fidelity coding strategy is adopted for regions of interest, and a high compression ratio coding strategy is adopted for regions of non-interest. The differentiated coding scheme includes quantization parameters, bitrate control mode, and frame type allocation rules.
[0079] Encoding priority parameters are determined based on encoding requirement indicators output by the network transmission adaptation model and are used to guide the allocation strategy of encoding resources for different video content. Content saliency analysis results are obtained by performing content saliency analysis on real-time video frame sequences, including the division of regions of interest (ROI) and non-ROI regions. Differentiated coding schemes are different encoding strategies formulated for ROI and non-ROI regions based on encoding priority parameters and content saliency analysis results. High-fidelity coding strategies aim to preserve as much detail and quality as possible in the video content during encoding, using parameters such as a smaller quantization step size and a higher bitrate. High compression ratio coding strategies compress video content during encoding by using parameters such as a larger quantization step size and a lower bitrate to reduce data volume. Quantization parameters are used to control quantization errors during encoding, bitrate control modes are used to control the video encoding bitrate, and frame type allocation rules are used to determine which frame type (e.g., I-frame, P-frame, B-frame) should be used for encoding video frames.
[0080] In some implementations, step S330 may include the following steps S331 to S336:
[0081] Step S331: Compare the encoding priority parameter with the preset priority threshold. When the encoding priority parameter is greater than the preset priority threshold, the corresponding area is determined to be a high-priority encoding area. When the encoding priority parameter is less than or equal to the preset priority threshold, the corresponding area is determined to be a low-priority encoding area. The preset priority threshold is determined through statistical analysis of historical output data of the network transmission adaptation model.
[0082] A preset priority threshold is a pre-defined value used to distinguish between high-priority and low-priority coding regions. It is obtained through statistical analysis of historical output data from the network transport adaptation model; for example, the average or median of the coding priority parameters from historical outputs can be calculated as the preset priority threshold. High-priority coding regions are those with coding priority parameters greater than the preset priority threshold. These regions are typically important areas in a video frame and require high-fidelity coding strategies. Low-priority coding regions are those with coding priority parameters less than or equal to the preset priority threshold. These regions are typically less important areas in a video frame and can be encoded using high-compression-ratio strategies.
[0083] Step S332: For the high-priority coding region, call the first quantization parameter set from the quantization parameter library. The first quantization parameter set is determined in the following way: based on the mapping relationship between the coding priority parameter and the quantization step size, when the coding priority parameter increases by one unit, the quantization step size decreases by a preset ratio, so that the quantization error of the high-priority coding region is controlled within the preset error range.
[0084] The quantization parameter library is a database storing different sets of quantization parameters, each corresponding to different encoding requirements and network states. The first quantization parameter set is used for high-priority encoding regions and is determined based on the mapping relationship between encoding priority parameters and quantization step size. The quantization step size is a parameter used to control quantization error during encoding; a smaller quantization step size results in smaller quantization error and higher encoding quality. The preset ratio is a pre-defined value used to determine the percentage decrease in the quantization step size when the encoding priority parameter increases. The preset error range is a pre-defined numerical range used to control the quantization error in high-priority encoding regions.
[0085] Step S333: For low-priority coding regions, call the second quantization parameter set from the quantization parameter library. The quantization step size of the second quantization parameter set is a preset multiple of the quantization step size of the first quantization parameter set. The preset multiple is dynamically adjusted according to the current availability of network transmission bandwidth. The lower the bandwidth availability, the larger the preset multiple.
[0086] The second quantization parameter set is used for low-priority encoding regions, and its quantization step size is a preset multiple of the quantization step size of the first quantization parameter set. This preset multiple is a dynamically adjusted value based on the current availability of network bandwidth; the lower the bandwidth availability, the larger the preset multiple. The current availability of network bandwidth refers to the bandwidth resources currently available for video transmission within the network.
[0087] Step S334: For high-priority coding regions, a content-based bitrate control mode is adopted, and bitrate resources are dynamically allocated by analyzing the texture complexity and motion intensity of the region.
[0088] Specifically, the entropy parameter in the texture features of the region can be extracted, and the entropy parameter can be compared with a preset texture threshold. When the entropy parameter is greater than the preset texture threshold, it is determined to be a region with high texture complexity. The preset texture threshold is determined by statistical analysis of the entropy values of texture-rich regions in historical video frames. The mean amplitude and density of motion vectors in the motion features of the region can be extracted, and the mean amplitude of motion vectors can be compared with a preset motion threshold, and the motion vector density can be compared with a preset density threshold. When the mean amplitude of motion vectors is greater than the preset motion threshold and the motion vector density is greater than the preset density threshold, it is determined to be a region with intense motion. The preset motion threshold and preset density threshold are determined by statistical analysis of the motion features of dynamic targets. The highest bitrate weight is assigned to sub-regions that simultaneously meet the criteria of high texture complexity and intense motion, the medium bitrate weight is assigned to sub-regions that only meet one of the criteria, and the basic bitrate weight is assigned to sub-regions that do not meet either criterion, ensuring that bitrate resources are tilted towards regions with rich details.
[0089] Step S335: For low-priority coding regions, adopt a total bitrate control mode, set the maximum bitrate limit for the region, and distribute the remaining bitrate resources evenly in the low-priority region through a global bitrate allocation algorithm to avoid the low-priority region occupying too much transmission resources.
[0090] Total bitrate control (TBI) mode sets a total bitrate cap when allocating bitrate resources, and then distributes the remaining bitrate resources across different regions. The maximum bitrate cap is a pre-set value used to limit bitrate usage in low-priority coding regions, preventing them from consuming excessive transmission resources. The global bitrate allocation algorithm is used to evenly distribute the remaining bitrate resources across low-priority regions.
[0091] In practice, the first step is to set a maximum bitrate cap for low-priority coding regions. This cap can be determined based on the current availability of network bandwidth and the bitrate requirements of high-priority coding regions. Then, a global bitrate allocation algorithm is used to evenly distribute the remaining bitrate resources across the low-priority regions. This algorithm can employ an average allocation method, dividing the remaining bitrate resources by the number of low-priority regions to obtain the average bitrate for each region.
[0092] Step S336: During the frame type allocation rule formulation process, calculate the area ratio of high-priority coding regions and low-priority coding regions in the video frame. Compare the area ratio of high-priority coding regions with a preset area ratio. When the area ratio of high-priority coding regions is greater than the preset area ratio, determine that the video frame adopts an independent coding frame type. The independent coding frame type completes encoding through its own pixel data and does not rely on the reference information of other video frames. Compare the area ratio of low-priority coding regions with a preset area ratio. When the area ratio of low-priority coding regions is greater than the preset area ratio, determine that the video frame adopts a bidirectional reference coding frame type. The bidirectional reference coding frame type achieves predictive encoding by referencing the encoding data of adjacent video frames. The preset area ratio is determined through a balance analysis of encoding efficiency and decoding complexity. Specifically, it is determined by constructing an objective function of encoding time and decoding time, and solving for the area ratio threshold that minimizes the objective function value as the preset area ratio.
[0093] When formulating frame type allocation rules, the first step is to accurately calculate the area ratio of high-priority and low-priority coded regions in a video frame. For a video frame, it can be divided into individual pixel units. The number of pixels contained in each high-priority and low-priority coded region can be counted, and then divided by the total number of pixels in the video frame to obtain the corresponding area ratio.
[0094] The preset area ratio is obtained through a balance analysis of encoding efficiency and decoding complexity. In the specific implementation, an objective function is constructed that considers both encoding and decoding time. For example, let the encoding time be T. encode Decoding takes T seconds. decode The objective function can be expressed as F=aT encode +bT decode Where a and b are weighting coefficients, adjusted according to the actual application scenario. Through extensive experiments and data analysis, the area proportion is continuously changed, the value of the objective function is calculated, and the area proportion threshold that minimizes the objective function value is determined. This threshold is then used as the preset area proportion.
[0095] Step S340: Generate video coding control instructions based on the differentiated coding scheme. The video coding control instructions include coding mode selection signals, parameter adjustment signals, and region division signals. The coding mode selection signal is used to specify the intra-frame or inter-frame coding mode. The parameter adjustment signal is used to dynamically set the quantization step size and bit rate control parameters. The region division signal is used to identify regions of interest and regions of non-interest.
[0096] Differential coding schemes determine different coding strategies for regions of interest (ROIs) and non-ROIs, including quantization parameters, rate control modes, and frame type allocation rules. Based on these strategies, video coding control instructions are generated. The coding mode selection signal is determined according to the frame type allocation rules. If the video frame uses an independent coding frame type, the coding mode selection signal specifies intra-frame coding mode; if it uses a bidirectional reference coding frame type, it specifies inter-frame coding mode. For example, for a video frame determined to use an independent coding frame type, the coding mode selection signal will instruct the encoder to use intra-frame coding mode, that is, to encode only the pixel data of the current frame.
[0097] The parameter adjustment signal is generated based on the quantization parameters and the bit rate control mode. For high-priority coding regions, the parameter adjustment signal sets a smaller quantization step size and a higher bit rate control parameter to ensure high-fidelity coding. For low-priority coding regions, the parameter adjustment signal sets a larger quantization step size and a lower bit rate control parameter to achieve high compression ratio coding.
[0098] Region segmentation signals are generated based on content saliency analysis results and are used to identify regions of interest (ROI) and non-ROI regions in a video frame. The encoder can apply different coding strategies to different regions based on the region segmentation signal. For example, after receiving the region segmentation signal, the encoder knows which regions are ROI and which are non-ROI, and thus applies a high-fidelity coding strategy to the ROI and a high compression ratio coding strategy to the non-ROI regions.
[0099] Step S350: Input the video encoding control command into the encoder to perform content-aware encoding conversion on the real-time video frame sequence. The encoder applies differentiated encoding strategies to different areas according to the control command to generate encoded video data units.
[0100] The encoder receives video encoding control commands and encodes the real-time video frame sequence according to these commands. Upon receiving the encoding mode selection signal, the encoder selects an appropriate encoding mode, such as intra-frame coding mode or inter-frame coding mode. For video frames using intra-frame coding mode, the encoder performs operations such as Discrete Cosine Transform (DCT) on the pixel data of the current frame to convert the pixel data into frequency domain coefficients, and then performs quantization and entropy coding. For video frames using inter-frame coding mode, the encoder first performs motion estimation and compensation, finds similar parts with the preceding and following frames, and then encodes the differences.
[0101] Upon receiving the parameter adjustment signal, the encoder dynamically sets the quantization step size and bit rate control parameters. For high-priority coding regions, the encoder uses a smaller quantization step size to reduce quantization errors and ensure coding quality. After receiving the region segmentation signal, the encoder employs different coding strategies for regions of interest (ROI) and non-ROI regions. For ROI, the encoder uses a high-fidelity coding strategy, such as using finer quantization and coding algorithms, to ensure image quality in that region. For non-ROI regions, the encoder uses a high compression ratio coding strategy, such as using coarser quantization and coding algorithms, to reduce data volume.
[0102] Step S360: Encapsulate the encoded video data units, add an encoding header containing encoding strategy information and region division information to obtain the encoding optimization stream. The data units in the encoding optimization stream are arranged in timestamp order.
[0103] The encapsulation process combines encoded video data units with a codec header to form a complete data packet. The codec header contains encoding strategy information and region segmentation information, which are crucial for decoding and processing at the receiving end. Encoding strategy information includes encoding mode, quantization parameters, and bitrate control mode, allowing the receiving end to select appropriate decoding algorithms and parameters. Region segmentation information identifies regions of interest (ROI) and non-ROI within the video frame, enabling the receiving end to apply different decoding strategies to different regions.
[0104] Step S400: During the transmission of the coded optimization stream, link status fluctuation information is obtained through the 5G network feedback channel. Based on the link status fluctuation information, the transmission strategy of the coded optimization stream is dynamically calibrated to obtain the calibrated transmission stream.
[0105] The 5G network feedback channel is a dedicated channel in 5G networks for transmitting link status information, providing real-time feedback of the network's current state to the transmitting end. Link status fluctuation information reflects the dynamic changes in the 5G network transmission link at different times, such as fluctuations in network bandwidth, changes in signal strength, and packet loss rates. By acquiring this link status fluctuation information, the transmitting end can promptly understand network changes and dynamically calibrate the transmission strategy for coded optimization streams.
[0106] Dynamic calibration of the transmission strategy adjusts parameters such as transmission bitrate, scheduling priority, and error correction redundancy of the encoded optimized stream based on link state fluctuation information to ensure efficient and stable video transmission under different network environments. The calibrated transmission stream is the video stream obtained after dynamic calibration of the transmission strategy, which can better adapt to network changes and improve video transmission quality and reliability.
[0107] In some implementations, step S400 may include the following steps S410 to S470:
[0108] Step S410: Periodically receive link status sampling data through the radio resource control interface of the 5G network. The link status sampling data includes the bandwidth fluctuation coefficient, delay jitter variance and packet loss density distribution of the transmission link in the continuous sampling period.
[0109] The bandwidth fluctuation coefficient is calculated by the degree of deviation between the current period bandwidth and the previous period bandwidth. The delay jitter variance reflects the degree of dispersion of transmission delay, and the packet loss density distribution characterizes the spatial clustering characteristics of packet loss events per unit time.
[0110] The radio resource control interface (RRI) of a 5G network is the interface between the 5G base station and terminal equipment used to control and manage radio resources. Through this interface, link status sampling data can be received periodically. The sampling period is a preset time interval, such as once per second. During each sampling period, the base station samples and measures various parameters of the transmission link.
[0111] Delay jitter variance is an indicator reflecting the dispersion of transmission delay, obtained through statistical analysis of transmission delay over multiple sampling periods. A larger delay jitter variance indicates more unstable transmission delay variations. Packet loss density distribution is an indicator characterizing the spatial clustering of packet loss events per unit time, obtained by statistically analyzing the number of packet losses at different locations within a unit of time. For example, within one second, packet loss is monitored at different locations in the network, the number of packet losses at each location is recorded, and then the distribution of these packet loss counts is analyzed. If packet loss is concentrated at certain defined locations, it indicates that the packet loss density distribution exhibits spatial clustering characteristics.
[0112] Step S420: Perform multi-dimensional correlation analysis on the link state sampling data, calculate the covariance value of bandwidth fluctuation coefficient and latency jitter variance, determine the coupling strength between the two, and analyze the correlation coefficient between packet loss density distribution and bandwidth fluctuation coefficient to construct a link state correlation matrix. The row vectors of the link state correlation matrix contain standardized feature values of bandwidth fluctuation, latency jitter and packet loss density.
[0113] Multidimensional correlation analysis analyzes and processes link state sampling data from multiple different angles and dimensions to study the interrelationships between different parameters. Covariance is an indicator that measures the linear relationship between two variables. By calculating the covariance between bandwidth fluctuation coefficient and latency jitter variance, the coupling strength between the two can be determined. If the covariance value is positive, it indicates that the bandwidth fluctuation coefficient and latency jitter variance are positively correlated, that is, the greater the bandwidth fluctuation, the greater the latency jitter; if the covariance value is negative, it indicates that the two are negatively correlated; if the covariance value is 0, it indicates that there is no linear relationship between the two.
[0114] The correlation coefficient is another indicator that measures the linear relationship between two variables. It is obtained by dividing the covariance by the product of the standard deviations of the two variables, and its value ranges from [-1, 1]. By analyzing the correlation coefficient between packet loss density distribution and bandwidth fluctuation coefficient, we can understand the degree of association between the two.
[0115] When constructing the link-state correlation matrix, the eigenvalues of bandwidth fluctuation, latency jitter, and packet loss density are first standardized to ensure they have the same scale and range. Standardization can be achieved using the Z-score standardization method, which involves subtracting the mean from each eigenvalue and then dividing by its standard deviation. The standardized eigenvalues are then used as row vectors in the link-state correlation matrix. The link-state correlation matrix visually demonstrates the relationship between bandwidth fluctuation, latency jitter, and packet loss density, providing crucial information for subsequent transmission strategy calibration.
[0116] Step S430: Generate a transmission strategy calibration parameter set based on the link state correlation matrix. The calibration parameter set includes a rate adjustment coefficient, a scheduling weight factor, and an error correction redundancy. The rate adjustment coefficient is dynamically calculated by the ratio of the bandwidth fluctuation coefficient to the preset baseline fluctuation value. The scheduling weight factor is adjusted inversely based on the square root of the delay jitter variance. The error correction redundancy is determined based on the entropy parameter of the packet loss density distribution. The higher the entropy parameter, the greater the redundancy.
[0117] The transmission strategy calibration parameter set is a set of parameters generated based on the link state correlation matrix, used to dynamically calibrate the transmission strategy of the coded optimization stream. The bitrate adjustment coefficient is a parameter used to adjust the transmission bitrate of the coded optimization stream, dynamically calculated as the ratio of the bandwidth fluctuation coefficient to a preset baseline fluctuation value. The scheduling weight factor is a parameter used to adjust the priority of the transmission scheduling queue, adjusted inversely based on the square root of the delay jitter variance. A larger square root value of the delay jitter variance indicates more unstable transmission delay changes; in this case, the scheduling weight factor will decrease, lowering the scheduling priority of the data packet.
[0118] Error correction redundancy is a parameter used to improve transmission reliability, determined based on the entropy parameter of the packet loss density distribution. The entropy parameter measures the uncertainty of the packet loss density distribution; a higher entropy parameter indicates a more random distribution of packet loss, requiring increased error correction redundancy to ensure reliable data transmission.
[0119] Step S440: Dynamically scale the transmission bitrate of each data unit in the encoded optimized stream according to the bitrate adjustment coefficient. The bitrate adjustment range is determined by multiplying the bitrate adjustment coefficient with the content importance score of the data unit. The content importance score is obtained by weighted fusion of the salience value of the corresponding video region of the data unit and the encoding priority parameter. The higher the content importance score, the smaller the bitrate adjustment range.
[0120] The bitrate adjustment factor is a parameter dynamically calculated based on bandwidth fluctuations, used to adjust the transmission bitrate of each data unit in the encoded optimized stream. The content importance score of a data unit is an indicator measuring its importance, obtained by weighted fusion of the saliency value of the corresponding video region and the encoding priority parameter. The saliency value is obtained through content saliency analysis, reflecting the visual importance of the video region; the encoding priority parameter is determined based on the encoding requirement index output by the network transmission adaptation model, reflecting the priority of encoding resource allocation for different video content.
[0121] In the actual adjustment process, each data unit in the encoded optimization stream is first mapped to a region to determine its spatial coordinate range within the video frame. Then, the saliency value within this coordinate range is extracted, along with the corresponding encoding priority parameter. Next, the bitrate adjustment magnitude is calculated, which is equal to the product of the bitrate adjustment coefficient and the content importance score. The higher the content importance score, the smaller the actual impact of the bitrate adjustment magnitude, ensuring the bitrate stability of highly important data units.
[0122] In some implementations, step S440 may include the following steps S441 to S446:
[0123] Step S441: Perform region mapping on the data units in the encoded optimization stream, determine the spatial coordinate range of each data unit in the video frame, extract the saliency value within the coordinate range, the saliency value comes from the saliency map in the content saliency analysis results, and obtain the encoding priority parameter corresponding to the data unit, the encoding priority parameter is output by the network transmission adaptation model.
[0124] In the encoded and optimized stream, data units are encoded and encapsulated data packets, each corresponding to a region within a video frame. Region mapping associates data units with spatial coordinate ranges within the video frame, determining the specific location of each data unit within the frame. In practice, the spatial coordinate range of a data unit within a video frame can be determined based on the region division information in the encoding header, as well as the encoding method and order of the data units.
[0125] The saliency map is obtained through content saliency analysis and records the saliency value of each pixel in a video frame. After determining the spatial coordinate range of the data unit, the saliency values within that coordinate range are extracted from the saliency map.
[0126] The encoding priority parameter is output by the network transport adaptation model and is determined based on the current network status and the importance of the video content. When obtaining the encoding priority parameter corresponding to a data unit, the corresponding encoding priority parameter can be found from the output of the network transport adaptation model based on the encoding strategy information and region division information of the data unit.
[0127] Step S442: The saliency value and the coding priority parameter are weighted and fused. During the fusion process, the weight coefficient of the saliency value is higher than the weight coefficient of the coding priority parameter. The sum of the two weight coefficients is 1, and the content importance score of the data unit is obtained. The value range of the content importance score is consistent with the value range of the coding priority parameter.
[0128] In the weighted fusion process, the weight coefficient of the saliency value is higher than that of the encoding priority parameter. This is because the saliency value directly reflects the visual importance of a video region and has a greater impact on the user's visual experience. To ensure that the range of the content importance score is consistent with the range of the encoding priority parameter, the result after weighted fusion can be normalized. For example, if the range of the encoding priority parameter is [0,1], and the result after weighted fusion exceeds this range, it can be mapped to the range [0,1].
[0129] Step S443: Calculate the bitrate adjustment range. The bitrate adjustment range is equal to the product of the bitrate adjustment coefficient and the content importance score. The value range of the bitrate adjustment coefficient is [0,1]. The higher the content importance score, the smaller the actual impact of the bitrate adjustment range.
[0130] The bitrate adjustment range is a parameter used to adjust the transmission bitrate of a data unit. It is calculated by multiplying the bitrate adjustment coefficient by the content importance score. When the bitrate adjustment coefficient is 0, it means that the transmission of that data unit needs to be stopped; when the bitrate adjustment coefficient is 1, it means that no bitrate adjustment is needed for the data unit. The higher the content importance score, the smaller the actual impact of the bitrate adjustment range.
[0131] Step S444: Dynamically scale the current transmission bitrate of the data unit based on the bitrate adjustment range. When the bitrate adjustment coefficient is greater than the preset adjustment threshold, reduce the current transmission bitrate by the bitrate adjustment range. The reduced bitrate value is not lower than the preset minimum bitrate threshold. The preset minimum bitrate threshold is determined based on the basic resolution requirements of the video frame.
[0132] The preset adjustment threshold is a pre-defined critical value for the bitrate adjustment coefficient, such as 0.5. When the bitrate adjustment coefficient is greater than the preset adjustment threshold, it indicates that the network bandwidth fluctuates significantly, requiring adjustment of the data unit's transmission bitrate. The current transmission bitrate of the data unit is dynamically scaled according to the bitrate adjustment magnitude, i.e., the adjusted bitrate = current transmission bitrate × (1 - bitrate adjustment magnitude). The preset minimum bitrate threshold is a lower limit bitrate value determined based on the basic resolution requirements of video frames. If the adjusted bitrate is lower than the preset minimum bitrate threshold, the bitrate is adjusted to the preset minimum bitrate threshold.
[0133] Step S445: Smoothly transition the scaled bitrate value by setting an upper limit for the bitrate change gradient between adjacent data units. The upper limit for the bitrate change gradient is determined by the ratio of the average bitrate of the encoded optimized stream to the frame rate.
[0134] Smooth transitions are implemented to prevent sudden bitrate changes from affecting video playback. This is achieved by setting a maximum bitrate change gradient between adjacent data units. The maximum bitrate change gradient is the maximum bitrate change between adjacent data units, determined by the ratio of the average bitrate to the frame rate of the encoded optimized stream. When smoothing the transition of scaled bitrate values, the bitrate difference between two adjacent data units is calculated. If the bitrate difference exceeds the maximum bitrate change gradient, the bitrate change is limited to within that range.
[0135] Step S446: Record the bitrate adjustment history of each data unit and generate a bitrate adjustment trajectory. The bitrate adjustment trajectory includes the data unit identifier, bitrate values before and after adjustment, adjustment timestamp, and content importance score, which can be used as historical data reference for subsequent transmission strategy optimization.
[0136] Bitrate adjustment history records information about the bitrate adjustment process for each data unit. By generating a bitrate adjustment trajectory, this information can be organized and saved. The bitrate adjustment trajectory includes the data unit identifier, bitrate values before and after adjustment, adjustment timestamp, and content importance score. The data unit identifier is a unique number used to identify each data unit; the bitrate values before and after adjustment record the transmission bitrate of the data unit before and after the adjustment; the adjustment timestamp records the specific time of the bitrate adjustment; and the content importance score records the importance of the data unit.
[0137] During actual recording, for each data unit's bitrate adjustment operation, relevant information is recorded in the bitrate adjustment trajectory. This trajectory can be stored in a database for historical data reference in subsequent transmission strategy optimization. By analyzing the bitrate adjustment trajectory, we can understand the bitrate adjustment of different data units under different network conditions, summarize the optimal transmission strategy, and improve the quality and efficiency of video transmission.
[0138] Step S450: Adjust the priority weight of the transmission scheduling queue in combination with the scheduling weight factor, and perform sliding window weighted average processing on the delay margin of the data packets in the scheduling queue. The delay margin is calculated by subtracting the current system time and the expected transmission delay from the playback deadline. The expected transmission delay is predicted based on the historical delay mean and the current jitter variance in the link state sampling data. The scheduling weight factor increases as the weighted average of the delay margin decreases.
[0139] The transmission scheduling queue is used to manage and schedule data packets to be transmitted. By adjusting the priority weights of data packets in the queue, high-priority data packets can be guaranteed to be transmitted first. The scheduling weight factor is a parameter that is adjusted inversely based on the square root of the delay jitter variance, reflecting the scheduling priority of the data packets. Delay margin is an indicator of the remaining transmission time of a data packet, calculated by subtracting the current system time and the estimated transmission delay from the playback deadline. The playback deadline is the latest time that the data packet must arrive at the receiving end and be played. The current system time is the actual current time, and the estimated transmission delay is predicted based on the historical delay mean and the current jitter variance in the link state sampling data. The delay margin of the data packets in the scheduling queue is processed by a sliding window weighted average, where the sliding window is a fixed-size time window. Within each window, the delay margin of the data packets is calculated by weighted average, and the weight coefficient can be set according to the importance of the data packets or other factors. The scheduling weight factor increases as the weighted average of the delay margin decreases. When the delay margin of a data packet is small, it indicates that it needs to be transmitted as soon as possible, and the scheduling weight factor will increase, raising the scheduling priority of that data packet.
[0140] In some implementations, step S450 may include the following steps S451 to S458:
[0141] Step S451: Parse the timestamp sequence of data packets in the coding optimization stream, extract the timestamp interval features of adjacent data packets, and construct a data packet transmission timing model based on the timestamp interval features. The timing model includes the sending interval distribution and arrival interval distribution of data packets. The sending interval distribution is determined by the frame rate parameter of the coding optimization stream, and the arrival interval distribution is obtained based on the historical arrival time statistics in the link state sampling data.
[0142] The data packets in the encoded optimization stream contain timestamp information. A timestamp is a time identifier that marks each data packet and is used to determine the order of the data packets and their playback time. Parsing the timestamp sequence involves extracting the timestamp information from the data packets and sorting and analyzing these timestamps.
[0143] The timestamp interval characteristic of adjacent data packets is the difference in timestamps between two adjacent data packets. By analyzing these interval characteristics, the transmission pattern of data packets can be understood. A data packet transmission timing model is a model used to describe the timing pattern of data packet transmission, including the transmission interval distribution and arrival interval distribution. The transmission interval distribution is the probability distribution of the transmission time intervals of data packets at the sending end, which can be determined by the frame rate parameter of the encoded optimization stream. The arrival interval distribution is the probability distribution of the arrival time intervals of data packets at the receiving end, obtained based on historical arrival time statistics from link state sampling data.
[0144] Step S452: Perform delay trend analysis on the time series model, use the exponential smoothing algorithm to predict the trend of historical transmission delay sequences, and generate a delay prediction curve for the future preset time period. The slope of the delay prediction curve represents the delay change trend. A positive slope indicates that the delay is increasing, and a negative slope indicates that the delay is decreasing.
[0145] Exponential smoothing is a method used for time series forecasting. It predicts future data values by weighting historical data. When performing latency trend analysis on a time series model, the first step is to obtain the historical transmission latency sequence, which is the transmission latency data of data packets within multiple past sampling periods.
[0146] The slope of the latency prediction curve represents the trend of latency change. A positive slope indicates that the latency is increasing, suggesting that network transmission latency may become increasingly larger. A negative slope indicates that the latency is decreasing, suggesting that network transmission latency may become increasingly smaller.
[0147] Step S453: Dynamically adjust the priority evaluation window length based on the slope characteristics of the delay prediction curve. The larger the absolute value of the slope, the smaller the window length, to ensure a rapid response to the delay change trend. The window length adjustment range is positively correlated with the delay jitter variance in the link state correlation matrix.
[0148] The priority evaluation window length is the size of the time window used to evaluate the priority of data packets. By dynamically adjusting the window length, the priority of data packets can be adjusted in a timely manner according to the latency change trend. The larger the absolute value of the slope of the latency prediction curve, the more drastic the latency change. In this case, the priority evaluation window length needs to be reduced to respond to latency changes more quickly.
[0149] The window length adjustment range is positively correlated with the delay jitter variance in the link state correlation matrix. The larger the delay jitter variance, the more unstable the transmission delay is, and the more frequently the priority evaluation window length needs to be adjusted.
[0150] Step S454: Within the priority evaluation window, perform a multi-dimensional evaluation of the latency sensitivity of data packets. The latency sensitivity evaluation indicators include timestamp urgency, content importance attenuation coefficient, and historical transmission success rate. Timestamp urgency is calculated by the difference between the current time and the playback deadline. The content importance attenuation coefficient increases as timestamp urgency decreases. The historical transmission success rate is determined based on the statistical analysis of historical transmission records of the same type of data packets.
[0151] Latency sensitivity is a metric that measures how sensitive a data packet is to transmission latency. Multi-dimensional evaluation provides a more comprehensive understanding of the importance and urgency of data packets. Within the priority evaluation window, the latency sensitivity of each data packet is assessed.
[0152] The timestamp urgency is calculated by the difference between the current time and the playback deadline; the smaller the difference, the more urgent the data packet. The content importance decay coefficient is a coefficient that increases as the timestamp urgency decreases, reflecting how the importance of the data packet content changes over time. The historical transmission success rate is determined based on statistical analysis of historical transmission records of the same type of data packets, reflecting the reliability of data packet transmission.
[0153] Step S455: Input the timestamp urgency, content importance decay coefficient, and historical transmission success rate into the priority evaluation function to generate a comprehensive priority score. The priority evaluation function determines the weight coefficient of each indicator through the analytic hierarchy process, with timestamp urgency having the highest weight and historical transmission success rate having the lowest weight.
[0154] The priority evaluation function is used to calculate the overall priority score of data packets. It takes timestamp urgency, content importance attenuation coefficient, and historical transmission success rate as inputs, and generates the overall priority score through a weighted summation. The Analytic Hierarchy Process (AHP) is used to determine the weight coefficients. By comparing multiple indicators pairwise, their relative importance is determined, thus obtaining the weight coefficients for each indicator. In this embodiment, timestamp urgency has the highest weight, and historical transmission success rate has the lowest weight.
[0155] Step S456: Determine the final priority weight based on the product of the comprehensive priority score and the scheduling weight factor. The scheduling weight factor is dynamically adjusted based on the latency jitter variance in the link state correlation matrix. The larger the latency jitter variance, the stronger the amplification effect of the scheduling weight factor on the comprehensive priority score.
[0156] The final priority weight is an indicator used to determine the priority of a data packet in the transmission scheduling queue. It is obtained by multiplying the comprehensive priority score by the scheduling weight factor. The scheduling weight factor is a parameter that is dynamically adjusted based on the delay jitter variance in the link state correlation matrix. The larger the delay jitter variance, the more unstable the transmission delay, and the stronger the amplification effect of the scheduling weight factor on the comprehensive priority score.
[0157] Step S457: Dynamically sort the transmission scheduling queue based on the final priority weight. Use the bubble sort algorithm to put data packets with high priority weights to the front. At the same time, add a priority stability mark to each data packet. The stability mark is determined by the rate of change of priority weight within three consecutive evaluation windows. Data packets with a rate of change less than a preset stability threshold are marked as stable priorities, which reduces the frequency of priority adjustment in subsequent scheduling.
[0158] Bubble sort is a simple sorting algorithm that moves data packets with higher priority weights to the front by repeatedly comparing and swapping adjacent elements. When dynamically sorting a transmission scheduling queue, the final priority weights of two adjacent data packets are first compared. If the priority weight of the preceding data packet is less than that of the following data packet, their positions are swapped. This process is repeated until the entire queue is sorted.
[0159] Priority stability is an indicator used to identify the stability of packet priority, determined by the rate of change of priority weights over three consecutive evaluation windows. The rate of change of priority weights is the percentage change in the priority weight of a packet between two adjacent evaluation windows. A preset stability threshold is a pre-defined critical value for the rate of change of priority weights, for example, 0.1. When the rate of change of priority weights is less than the preset stability threshold, the packet is marked as having a stable priority. For packets with stable priorities, the frequency of priority adjustments is reduced in subsequent scheduling to decrease the system's computational overhead.
[0160] Step S458: Generate a priority adjustment instruction. The instruction includes a data packet identifier, final priority weight, stability flag, and window length parameter. It is sent to the transmission scheduling module through the control channel of the 5G network. The transmission scheduling module reorganizes the transmission queue according to the instruction to ensure that high-priority data packets occupy transmission resources first.
[0161] The priority adjustment instruction is used to notify the transmission scheduling module to adjust the priority of the transmission queue. It includes a packet identifier, a final priority weight, a stability flag, and a window length parameter. The packet identifier is a number used to uniquely identify each packet. The final priority weight is a priority index determined by multiplying the comprehensive priority score by the scheduling weight factor. The stability flag indicates the stability of the packet priority. The window length parameter is the length of the priority evaluation window.
[0162] After generating the priority adjustment command, it is sent to the transmission scheduling module via the 5G network's control channel. The control channel is a dedicated channel in the 5G network for transmitting control information, offering high reliability and real-time performance. Upon receiving the priority adjustment command, the transmission scheduling module reorganizes the transmission queue according to the information in the command. Data packets with higher priority weights are placed at the front, ensuring they have priority access to transmission resources. For example, based on the final priority weight in the command, the highest priority data packet is placed at the front of the transmission queue, and so on, ensuring that data packets are transmitted in priority order.
[0163] Step S460: Dynamically configure forward error correction coding parameters based on error correction redundancy, divide the data units in the coding optimization stream into check blocks, and determine the size of the check block based on the content complexity of the data unit. The content complexity is calculated by combining the texture entropy value and the average value of the motion vector amplitude of the data unit. The higher the content complexity, the smaller the check block. Then, generate check data based on error correction redundancy and embed it into the data unit.
[0164] Forward error correction coding is a coding method used to improve the reliability of data transmission. By adding redundant information to the original data, the receiving end can recover the complete data even when it receives only partial data. The error correction redundancy is determined based on the entropy parameter of the packet loss density distribution, representing the proportion of redundant information that needs to be added.
[0165] When configuring forward error correction coding parameters, the specific coding parameters, such as the coding rate and generator polynomial, are first determined based on the error correction redundancy. Then, the data units in the coding optimization stream are divided into check blocks. A check block divides the data unit into several smaller blocks, each of which undergoes independent check and coding. The size of the check block is determined based on the content complexity of the data unit, which is calculated by combining the texture entropy value and the average magnitude of the motion vector of the data unit.
[0166] Texture entropy is a metric that measures the texture complexity of a data unit, obtained by analyzing the texture features of the data unit. The mean magnitude of motion vectors is a metric that measures the intensity of motion of objects within a data unit, obtained by analyzing the motion characteristics of the data unit.
[0167] The higher the content complexity, the smaller the check block should be. This is because complex data units are more prone to errors, and smaller check blocks can improve the accuracy of error correction.
[0168] After dividing the check blocks, check data is generated based on the error correction redundancy. The check data is redundant information obtained by applying a specific encoding algorithm (such as RS encoding, convolutional encoding, etc.) to the check blocks. Finally, the generated check data is embedded into the data units, ensuring that each data unit contains both the original data and the check data, thus improving the reliability of data transmission.
[0169] Step S470: The dynamically scaled bitrate data, the priority-adjusted scheduling queue, and the data units with embedded verification data are reassembled for transmission resources to generate a calibration transport stream. Each data unit in the calibration transport stream carries a transmission strategy calibration flag to instruct the receiving end to apply the corresponding decoding compensation strategy.
[0170] Transmission resource reorganization involves recombining and reorganizing various data that have undergone rate adjustment, priority adjustment, and error correction coding to form a complete transmission stream. The dynamically scaled rate data is obtained by dynamically scaling the transmission rate of each data unit in the coded optimization stream according to the rate adjustment coefficient. The priority-adjusted scheduling queue is obtained by adjusting the transmission scheduling queue according to the scheduling weight factor and the final priority weight. The data unit with embedded check data is a data unit that has undergone forward error correction coding and has embedded check data.
[0171] During the transmission resource reorganization process, the dynamically scaled bitrate data is first associated with the priority-adjusted scheduling queue to ensure that the transmission bitrate and priority information of each data packet are consistent. Then, the data units with embedded checksum data are arranged in the adjusted priority order to form an ordered data stream.
[0172] Each data unit carries a transmission strategy calibration flag, which includes encoding strategy information, region division information, bitrate adjustment information, priority adjustment information, etc., to instruct the receiving end to apply the corresponding decoding compensation strategy. For example, the transmission strategy calibration flag can use a set encoding format, with each field representing different information. After receiving the calibration transport stream, the receiving end selects an appropriate decoding algorithm and parameters based on the information in the transmission strategy calibration flag to decode and compensate the data, recovering the original video data. In this way, a calibration transport stream is generated, improving the quality and reliability of video transmission.
[0173] Step S500: Decode and align the calibration transport stream to generate a visual call output sequence that is synchronized with the original video stream unit time, and push it to the receiving end presentation device.
[0174] Decoding timing alignment is performed to ensure that video frames in the calibration transport stream are decoded and played in the correct time order, thereby generating a visual call output sequence that is time-synchronized with the original video stream units. During transmission, due to factors such as network latency and packet loss, the arrival time of video frames may deviate, requiring timing alignment to correct these deviations.
[0175] In some implementations, step S500 may include the following steps S510 to S570:
[0176] Step S510: Receive the calibration transport stream, extract the transmission strategy calibration mark and compressed data segment of each data unit, determine the strategy index through the mark field, and call the corresponding decoding compensation strategy. The decoding compensation strategy includes error hiding algorithm, motion compensation mode and texture restoration coefficient.
[0177] The calibration transport stream is a video data stream obtained after dynamic calibration using a transmission strategy. It contains encoded video data units and transmission strategy calibration markers. Upon receiving the calibration transport stream, each data unit is first parsed to extract its transmission strategy calibration marker and compressed data segment. The transmission strategy calibration marker includes encoding strategy information, region division information, bitrate adjustment information, priority adjustment information, etc. The strategy index can be determined through the marker field.
[0178] A policy index is a unique identifier for a decoding compensation policy. Based on the policy index, the corresponding decoding compensation policy can be retrieved from the decoding compensation policy library. The decoding compensation policy library is a database storing various decoding compensation policies. Each policy includes information such as error concealment algorithms, motion compensation modes, and texture restoration coefficients. Error concealment algorithms are used to handle errors and packet loss that occur during transmission, such as block-based error concealment algorithms and motion-based error concealment algorithms. Motion compensation modes are used for motion estimation and compensation during inter-frame coding, such as unidirectional motion compensation and bidirectional motion compensation. Texture restoration coefficients are parameters used to adjust the texture sharpness of the decoded video frames and can be dynamically adjusted based on information in the transmission policy calibration markers.
[0179] Step S520: Perform entropy decoding on the compressed data segment, switch the context model according to the encoding type of the data unit, and convert the binary code stream into a quantization coefficient matrix. The encoding type is determined by the start code identifier of the compressed data segment.
[0180] Entropy decoding is a decoding method used to restore compressed binary bitstreams to their original data. It involves decoding and parsing the bitstream to obtain a quantization coefficient matrix. This matrix, formed by the encoder's quantization process, contains frequency domain information about the video frames.
[0181] During entropy decoding, the context model is first switched based on the encoding type of the data unit. The encoding type can be determined by the start code identifier of the compressed data segment. Different encoding types (such as I-frames, P-frames, and B-frames) require different context models for decoding. The context model is a model used to describe the context information during the encoding process, including information such as the encoding probability distribution and symbol table. Then, entropy decoding is performed on the compressed data segment according to the context model, converting the binary bitstream into a quantization coefficient matrix. Specific algorithms for entropy decoding include, for example, Huffman coding and arithmetic coding. For instance, using Huffman coding for entropy decoding, the binary bitstream is decoded according to the Huffman tree to obtain the quantization coefficient matrix.
[0182] Step S530: Perform inverse quantization and inverse transform processing on the quantization coefficient matrix. The inverse quantization step size in the decoding compensation strategy is adopted. The inverse quantization step size is corrected by the quantization adjustment coefficient in the transmission strategy calibration mark. The transform kernel function is selected according to the intra-frame coding mode or inter-frame coding mode for inverse transform.
[0183] Dequantization is the process of restoring the quantization coefficient matrix to the original frequency domain coefficient matrix, achieved by multiplying by the dequantization step size. The dequantization step size is determined based on information from the decoding compensation strategy and corrected by the quantization adjustment coefficient in the transmission strategy calibration mark. The quantization adjustment coefficient is determined based on the bandwidth fluctuation coefficient in the link state fluctuation information, reflecting the impact of network bandwidth changes on the quantization step size. For example, when the network bandwidth is high, the quantization adjustment coefficient is large, and the dequantization step size increases accordingly; when the network bandwidth is low, the quantization adjustment coefficient is small, and the dequantization step size decreases accordingly. Inverse transform is the process of converting the frequency domain coefficient matrix into spatial domain pixel data. The inverse transform is performed by selecting a transform kernel function based on the intra-frame coding mode or inter-frame coding mode. In intra-frame coding mode, the inverse transform of the integer discrete cosine transform can be used; in inter-frame coding mode, the inverse transform of the integer discrete sine transform is typically used.
[0184] In some implementations, step S530 may include the following steps S531 to S535:
[0185] Step S531: Extract the basic inverse quantization step size from the decoding compensation strategy, and correct the inverse quantization step size by using the quantization adjustment coefficient in the transmission strategy calibration mark. The quantization adjustment coefficient is determined based on the bandwidth fluctuation coefficient in the link status fluctuation information.
[0186] The decoding compensation strategy includes a basic inverse quantization step size, which is determined based on the quantization parameters during encoding. The quantization adjustment coefficient in the transmission strategy calibration flag is determined based on the bandwidth fluctuation coefficient in the link state fluctuation information and is used to correct the basic inverse quantization step size.
[0187] In practice, the basic inverse quantization step size is first extracted from the decoding compensation strategy, and then the quantization adjustment coefficients are obtained from the transmission strategy calibration flags. The quantization adjustment coefficients can be determined based on the bandwidth fluctuation coefficient through a certain mapping relationship (such as a linear mapping). In this way, the inverse quantization step size is dynamically adjusted according to changes in network bandwidth, thereby improving the decoding quality.
[0188] Step S532: Perform dequantization calculation on each coefficient in the quantization coefficient matrix. Multiply the quantization coefficient by the corrected dequantization step size to obtain the frequency domain coefficient matrix. The dimension of the frequency domain coefficient matrix is consistent with the transform block size at the encoding end.
[0189] After obtaining the corrected inverse quantization step size, inverse quantization is performed on each coefficient in the quantization coefficient matrix. Inverse quantization involves multiplying the quantization coefficient by the corrected inverse quantization step size to obtain the frequency domain coefficient matrix. The dimension of the frequency domain coefficient matrix is consistent with the transform block size at the encoding end; this is to ensure the correctness of the inverse transform.
[0190] Step S533: Select the transform kernel function according to the coding mode of the data unit. In the intra-frame coding mode, the inverse transform of the integer discrete cosine transform is used, and in the inter-frame coding mode, the inverse transform of the integer discrete sine transform is used to perform the inverse transform on the frequency domain coefficient matrix.
[0191] There are two coding modes for data units: intra-frame coding mode and inter-frame coding mode. Different coding modes require different transform kernel functions for inverse transform. In intra-frame coding mode, the inverse transform of the Integer Discrete Cosine Transform (IDCT) is usually used to convert the frequency domain coefficient matrix into spatial domain pixel data. In inter-frame coding mode, the inverse transform of the Integer Discrete Sine Transform (IDST) can be used, which also converts the frequency domain coefficient matrix into spatial domain pixel data.
[0192] During the inverse transform, the transform kernel function is first determined based on the coding mode of the data unit. For example, for a data unit using an intra-frame coding mode, the inverse integer discrete cosine transform is chosen as the transform kernel function. Then, the frequency domain coefficient matrix is input into the transform kernel function for inverse transform calculation.
[0193] Step S534: Perform dynamic range adjustment on the spatial domain pixel data after inverse transformation. Map the pixel values to the color gamut range of the display device through linear mapping. The mapping coefficient is dynamically adjusted according to the average value of the luminance component of the video frame.
[0194] The pixel value range of the spatial domain pixel data after inverse transformation may not be consistent with the color gamut range of the display device, thus requiring dynamic range adjustment. Linear mapping is one adjustment method that maps pixel values linearly from the original range to the color gamut range of the display device.
[0195] The mapping coefficient is dynamically adjusted based on the average luminance component of the video frame. The average luminance component can be calculated by converting the video frame to a grayscale image. For example, for a color video frame, it is first converted to a grayscale image, and then the average luminance value of all pixels in the grayscale image is calculated as the average luminance component. When the average luminance component is high, it indicates that the video frame is generally bright; in this case, the mapping coefficient can be appropriately increased to make the pixel values appear brighter on the display device. Conversely, when the average luminance component is low, it indicates that the video frame is generally dark; in this case, the mapping coefficient can be appropriately decreased to make the pixel values appear darker on the display device. Through this dynamic adjustment, the video can achieve a good display effect on the display device under different brightness conditions.
[0196] Step S535: Store the adjusted spatial domain pixel data into the decoded pixel cache, manage the cache by frame number and macroblock position index, and use the least recently used algorithm to evict reference frame data whose content importance score is lower than the score threshold.
[0197] The decoded pixel buffer is a buffer used to store decoded spatial domain pixel data, improving decoding and playback efficiency. When storing adjusted spatial domain pixel data into the decoded pixel buffer, it is managed using indexes based on frame sequence number and macroblock position. The frame sequence number distinguishes different video frames, and the macroblock position distinguishes different macroblocks within the same video frame. For example, a video frame is divided into multiple macroblocks, each with a unique position index. Using the frame sequence number and macroblock position index, pixel data in the buffer can be quickly found and accessed.
[0198] The Least Recently Used (LRU) algorithm is a cache eviction algorithm that determines which data to evict based on its usage frequency. In the decoded pixel cache, reference frame data with a content importance score below a certain threshold is evicted using the LRU algorithm. The content importance score is obtained by performing content saliency analysis and encoding priority evaluation on video frames, reflecting the importance of each frame. The score threshold is a pre-set value used to determine whether reference frame data should be retained.
[0199] Step S540: Predictive compensation is performed on the inverse-transformed pixel data through motion compensation mode. In inter-frame coding mode, the forward or backward reference frame in the reference frame buffer is called, and pixel block matching is performed based on the motion vector. During the matching process, a sub-pixel interpolation algorithm is used to improve accuracy.
[0200] Motion compensation mode is used for motion estimation and compensation during inter-frame coding, reducing redundant information between video frames and improving coding efficiency. When performing prediction compensation on the pixel data after inverse transform, for video frames in inter-frame coding mode, it is necessary to call the forward or backward reference frame in the reference frame buffer. The forward reference frame is the video frame preceding the current video frame, and the backward reference frame is the video frame following the current video frame.
[0201] Pixel block matching is performed based on motion vectors, which are calculated during the encoding process and represent the direction and magnitude of a pixel block's displacement in adjacent video frames. Sub-pixel interpolation algorithms are used to improve accuracy during matching. Sub-pixel interpolation algorithms are used to estimate pixel values at the sub-pixel level, improving the accuracy of motion compensation. For example, feasible sub-pixel interpolation algorithms include bilinear interpolation and bicubic interpolation.
[0202] Step S550: Perform loop filtering on the compensated pixel data. The filtering intensity is dynamically adjusted according to the texture restoration coefficient. The higher the texture restoration coefficient, the larger the filter kernel size, thus eliminating block effect and ringing effect.
[0203] Loop filtering is a post-processing operation performed on pixel data during decoding. It can eliminate block artifacts and ringing artifacts generated during encoding, thereby improving the visual quality of the video. Block artifacts occur during encoding due to the use of block coding, resulting in obvious boundaries between adjacent blocks. Ringing artifacts occur during encoding because high-frequency components of the signal are truncated, causing oscillations at the edges of the signal.
[0204] The filtering intensity is dynamically adjusted based on the texture restoration coefficients. The texture restoration coefficients are a parameter in the decoding compensation strategy, reflecting the requirements for texture restoration of video frames. A higher texture restoration coefficient indicates a higher requirement for texture restoration, requiring a larger filter kernel size. The filter kernel size is the size of the loop filter, determining the filtering range. By dynamically adjusting the filtering intensity, blockiness and ringing artifacts can be eliminated while preserving as much texture information as possible from the video frames, thus improving video quality.
[0205] Step S560: Extract the timestamp of the video frame and calculate the deviation value with the system reference clock. When the deviation value is positive, generate a transition compensation frame by fusing the motion vectors of adjacent frames. When the deviation value is negative, filter out non-critical frames and discard them. Non-critical frames are determined by frame type identifier and content importance score.
[0206] A video frame's timestamp is a time identifier added to each video frame during the encoding process to determine the playback order and timing of the video frames. The system reference clock is a precise clock used to provide a standard time reference. The extracted video frame timestamp is compared to the system reference clock, and the deviation value between the two is calculated. The deviation value is the difference between the video frame timestamp and the system reference clock.
[0207] When the deviation value is positive, it indicates that the playback time of the video frame is later than the system reference clock time. In this case, a transition compensation frame is generated by fusing the motion vectors of adjacent frames. Adjacent frames are the frames before and after the current video frame. Motion vector fusion involves weighted averaging of the motion vectors of these adjacent frames to obtain the motion vector of the transition compensation frame. The pixel data of the transition compensation frame is then generated based on the motion vector of the transition compensation frame and inserted into the current video frame to compensate for the time deviation.
[0208] When the deviation value is negative, it indicates that the playback time of the video frame is earlier than the system reference clock. In this case, non-critical frames are selected and discarded. Non-critical frames are determined by their frame type identifier and content importance score. The frame type identifier distinguishes between critical frames (such as I-frames) and non-critical frames (such as P-frames and B-frames), while the content importance score reflects the importance of the video frame. For example, video frames with a non-critical frame type identifier and a content importance score below a certain threshold can be marked as discardable frames. In this way, the playback time of the video can be synchronized with the system reference clock, improving the playback quality.
[0209] In some implementations, step S560 may include the following steps S561 to S566:
[0210] Step S561: Extract the timestamp field of the video frame and align it with the system reference clock. Dynamically correct the timestamp offset by calibrating the time synchronization packet in the transport stream. Calculate the timestamp deviation value, which is the difference between the video frame timestamp and the system reference clock.
[0211] The timestamp field of a video frame is added to the video frame during encoding to identify its playback time. After extracting the timestamp field, it needs to be aligned with the system reference clock. The time synchronization packet in the calibration transport stream is a special data packet containing precise time information used to calibrate the timestamps of the video frames. By receiving the time synchronization packet, the timestamp offset can be dynamically corrected.
[0212] Step S562: Construct a sliding window sequence of deviation values, calculate the degree of temporal fluctuation by the root mean square of the deviation values within the window, and dynamically adjust the window length according to the video frame rate.
[0213] A sliding window sequence of deviation values is a fixed-size window used to store timestamp deviation values over a recent period. By calculating the root mean square (RMS) of the deviation values within the window, the degree of time series volatility can be obtained. The RMS is a statistical indicator that reflects the fluctuation of data.
[0214] The window length is dynamically adjusted based on the video frame rate. The video frame rate is the number of frames played per second, which determines the smoothness of the video. When the video frame rate is high, it means the video changes quickly, and the window length can be set smaller to more promptly reflect changes in timestamp deviation values. When the video frame rate is low, it means the video changes slowly, and the window length can be set larger to reduce statistical errors.
[0215] Step S563: When the deviation value is positive and the temporal fluctuation level is lower than the preset fluctuation threshold, extract the adjacent reference frames before and after, calculate the motion vector field of the two frames and dynamically allocate the fusion weight according to the deviation value, and generate the pixel data of the transition compensation frame based on the fused motion vector field.
[0216] When the deviation value is positive and the timing fluctuation is below the preset fluctuation threshold, it indicates that the playback time of the video frame is later than the system reference clock time, and the fluctuation of the timestamp deviation value is small. In this case, a transition compensation frame can be generated by motion vector fusion of adjacent frames. The adjacent reference frames are the frame before and the frame after the current video frame.
[0217] First, the motion vector fields of adjacent reference frames are calculated. A motion vector field is the collection of motion vectors for each pixel in a video frame, which can be calculated using motion estimation algorithms (such as optical flow estimation). Then, fusion weights are dynamically assigned based on the magnitude of the deviation. For example, a smaller deviation indicates a smaller temporal deviation, allowing for a larger weighting of the motion vectors from the previous frame; conversely, a larger deviation indicates a larger temporal deviation, allowing for a larger weighting of the motion vectors from the subsequent frame.
[0218] Pixel data for transition compensation frames is generated based on the fused motion vector field. This can be achieved by interpolating and predicting pixel data from adjacent reference frames using the fused motion vector field. For example, a bilinear interpolation algorithm can be used to interpolate pixel data from adjacent reference frames using the fused motion vector field to generate the transition compensation frame's pixel data. This method compensates for temporal discrepancies in video frames, ensuring continuous video playback.
[0219] Step S564: When the deviation value is negative and the temporal fluctuation exceeds the preset fluctuation threshold, filter video frames with the frame type identified as bidirectional reference coding frames, calculate the content importance score, and determine the score by multiplying the proportion of the intraframe with the area of intense motion and the texture complexity. Frames with scores below the threshold are marked as frames to be discarded.
[0220] When the deviation value is negative and the timing fluctuation exceeds the preset fluctuation threshold, it indicates that the playback time of the video frame is earlier than the system reference clock time, and the timestamp deviation value fluctuates significantly. In this case, non-critical frames need to be filtered and discarded. First, video frames with the frame type identifier of bidirectional reference coded frames are filtered. Bidirectional reference coded frames are usually not critical frames and have a relatively small impact on the overall playback of the video.
[0221] Calculate the content importance score for these bidirectional reference coded frames. The content importance score is determined by multiplying the proportion of intra-frame regions with intense motion by the texture complexity. The proportion of intra-frame regions with intense motion can be obtained by performing motion analysis on the video frame, for example, by using an optical flow estimation algorithm to calculate the proportion of regions in the video frame whose motion vector magnitude is greater than a preset threshold. Texture complexity can be obtained by calculating the texture entropy value of the video frame.
[0222] Frames with scores below a threshold are marked as frames to be discarded. The threshold is a pre-set value used to determine whether to discard the frame. In this way, the number of video frames can be reduced while maintaining basic video playback quality, and the video playback time can be synchronized with the system reference clock.
[0223] Step S565: Generate a transition pixel band for the frame to be discarded. This is done by linear interpolation of the boundary pixels of adjacent frames. The width of the transition pixel band is determined based on the intensity of the motion of the frame to be discarded. Remove the frame to be discarded and embed the transition pixel band.
[0224] Once the frames to be discarded are determined, a transition pixel band needs to be generated to avoid noticeable jumps during video playback. The transition pixel band is generated through linear interpolation of the boundary pixels of adjacent frames. The width of the transition pixel band is determined based on the motion intensity of the frame to be discarded. Motion intensity can be measured by the average magnitude or density of the motion vectors of the frame to be discarded. Higher motion intensity indicates faster frame changes, allowing for a wider transition pixel band width; lower motion intensity indicates slower frame changes, allowing for a narrower transition pixel band width. The frame to be discarded is then removed from the video sequence, and the generated transition pixel band is inserted into its position to make the video playback smoother.
[0225] Step S566: Recalculate the deviation value for the processed video frame sequence, and correct the fusion weight or discard threshold through a feedback adjustment mechanism. The feedback adjustment coefficient is dynamically determined according to the degree of temporal fluctuation, so that the deviation value converges to the preset synchronization range.
[0226] After generating transition compensation frames and processing frames to be discarded, the deviation value needs to be recalculated for the processed video frame sequence. This is done by extracting the timestamps of the video frames again and comparing them with the system reference clock to calculate the new deviation value. Then, a feedback adjustment mechanism is used to correct the fusion weights or the discard threshold.
[0227] The feedback adjustment coefficient is dynamically determined based on the degree of timing fluctuation. When the timing fluctuation is large, it indicates that the timestamp deviation value is unstable. In this case, the feedback adjustment coefficient can be set larger to adjust the fusion weights or discard thresholds more quickly. When the timing fluctuation is small, it indicates that the timestamp deviation value is relatively stable. In this case, the feedback adjustment coefficient can be set smaller to avoid over-adjustment. By continuously recalculating the deviation value and adjusting the fusion weights or discard thresholds, the deviation value is brought closer to the preset synchronization range. The preset synchronization range is a pre-defined range of deviation values, such as [-0.1, 0.1] seconds. When the deviation value is within this range, the playback time of the video frames is considered to be basically synchronized with the system reference clock. Through this feedback adjustment mechanism, the consistency between the video playback timing and the system reference clock can be guaranteed, improving the video playback quality.
[0228] Step S570: The processed video frames are spliced together in timestamp order to generate a visual call output sequence. The time synchronization accuracy is ensured by using a time sequence check code, which is calculated by the difference between the timestamps of adjacent frames.
[0229] The processed video frames are obtained after a series of processes such as decoding and timing alignment. These video frames are then spliced together in timestamp order. The timestamp is an important identifier of the video frame, determining its playback order.
[0230] Timing synchronization accuracy is ensured through a timing checksum. The timing checksum is calculated by subtracting the timestamps of adjacent frames. Upon receiving the video call output sequence, the receiving end calculates the timing checksums of adjacent frames and compares them with the timing checksums generated by the sending end. If they match, the timing synchronization accuracy is high; if they do not match, there may be a time deviation, requiring further adjustment. This method guarantees the timing synchronization accuracy of the video call output sequence during transmission and playback, providing high-quality video call video to the receiving end's display device. Finally, the generated video call output sequence is pushed to the receiving end's display device, such as a mobile phone screen or monitor, enabling both parties to conduct a clear and smooth video call.
[0231] It is understood that the various algorithms involved in the above descriptions of the embodiments of the present invention can all be obtained from relevant content in the prior art. To save space, they will not be elaborated on in the embodiments of the present invention. In addition, those skilled in the art can supplement the details based on common knowledge in the art when implementing the solutions of the present invention. For example, they can use normalization to eliminate dimensional conflicts before feature fusion, use interpolation to eliminate dimensional differences, reasonably set thresholds based on historical data, experience or business scenario requirements, train the model based on a general model training method, set the number of layers in the model structure based on actual needs, select activation functions, etc. The present invention will not provide redundant descriptions of overly detailed implementation processes here.
[0232] Please see Figure 3 This is a schematic diagram of the structure of a computer system provided in an embodiment of the present invention. Figure 3 As shown, the computer system 10 described above may include: a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the computer system 10 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to implement communication between these components. The user interface 1003 may optionally include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. The memory 1005 may also optionally be at least one storage device located remotely from the aforementioned processor 1001. Figure 3 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application.
[0233] exist Figure 3 In the computer system 10 shown, the network interface 1004 provides network communication functions; the user interface 1003 is mainly used to provide an input interface; and the processor 1001 can be used to call the device control application stored in the memory 1005 to implement the methods provided in the above embodiments.
Claims
1. A method for processing visual call information based on 5G, characterized in that, The method includes: The system continuously acquires real-time video frame sequences and 5G network environment perception data in a visual call scenario. The real-time video frame sequences are arranged in time stamp order to form video stream units, and the 5G network environment perception data reflects the dynamic change characteristics of the current transmission link. The 5G network environment perception data undergoes multi-dimensional state mapping processing to construct a network transmission adaptation model. Specifically, this includes: collecting wireless resource occupancy and signal quality change trends from the 5G network environment perception data; establishing a network state time-series database, which stores network characteristic variables according to a sampling period; performing spatiotemporal correlation analysis on the data in the network state time-series database to identify periodic patterns and sudden disturbance characteristics of network state fluctuations, generating a network state feature map; constructing a state transition probability matrix based on the network state feature map, which describes the migration probability and duration distribution between different network states; associating the state transition probability matrix with preset video coding requirement indicators to generate the basic framework of the network transmission adaptation model, which includes a network state input layer, a feature mapping layer, and a coding requirement output layer; dynamically adjusting the basic framework using historical transmission data to optimize the mapping relationship between network state and coding requirements, enabling the network transmission adaptation model to reflect the video coding adaptation requirements under the current network environment in real time. The network transmission adaptation model characterizes the dynamic correlation between network state and video coding requirements. Based on the network transmission adaptation model, video encoding control instructions are generated to perform content-aware encoding conversion on the real-time video frame sequence, outputting an optimized encoding stream. Specifically, this includes: parsing the encoding requirement indicators output by the network transmission adaptation model to determine video encoding priority parameters under the current network condition; these priority parameters guide the allocation strategy of encoding resources for different video content; performing content saliency analysis on the real-time video frame sequence to identify regions of interest (ROI) and non-ROI regions in the video frames; ROI regions are determined using a visual attention model, and non-ROI regions are areas in the video frames with saliency values below a threshold; and generating a differentiated encoding scheme based on the encoding priority parameters and the content saliency analysis results, employing a high-fidelity encoding strategy for ROI regions and a high compression ratio encoding strategy for non-ROI regions. The differentiated coding scheme includes quantization parameters, bitrate control mode, and frame type allocation rules. Video coding control instructions are generated based on the differentiated coding scheme. These instructions include coding mode selection signals, parameter adjustment signals, and region division signals. The coding mode selection signal specifies the intra-frame or inter-frame coding mode, and the parameter adjustment signal dynamically sets the quantization step size and bitrate control parameters. The video coding control instructions are input into the encoder to perform content-aware coding conversion on the real-time video frame sequence. The encoder applies differentiated coding strategies to different regions according to the control instructions, generating encoded video data units. The encoded video data units are encapsulated, and a coding header containing coding strategy information and region division information is added to obtain an optimized coding stream. The data units in the optimized coding stream are arranged in timestamp order. During the transmission of the coded optimization stream, link state fluctuation information is obtained through the 5G network feedback channel. Based on this link state fluctuation information, the transmission strategy of the coded optimization stream is dynamically calibrated to obtain a calibrated transmission stream. Specifically, this includes: periodically receiving link state sampling data through the 5G network's radio resource control interface. This link state sampling data includes the bandwidth fluctuation coefficient, latency jitter variance, and packet loss density distribution of the transmission link within a continuous sampling period. Multi-dimensional correlation analysis is performed on the link state sampling data to calculate the covariance between the bandwidth fluctuation coefficient and the latency jitter variance, determining their coupling strength. Simultaneously, the correlation coefficient between the packet loss density distribution and the bandwidth fluctuation coefficient is analyzed to construct a link state correlation matrix. The row vectors of this correlation matrix contain standardized eigenvalues of bandwidth fluctuation, latency jitter, and packet loss density. A transmission strategy calibration parameter set is generated based on the link state correlation matrix. This parameter set includes a rate adjustment coefficient, a scheduling weight factor, and error correction redundancy. The transmission rate of each data unit in the coded optimization stream is dynamically scaled according to the rate adjustment coefficient, and the rate adjustment coefficient is correlated with the content importance of the data unit. The product of scores determines the bitrate adjustment range; the priority weight of the transmission scheduling queue is adjusted in conjunction with the scheduling weight factor, and the delay margin of the data packets in the scheduling queue is processed by a sliding window weighted average. The delay margin is calculated by subtracting the current system time and the expected transmission delay from the playback deadline. The expected transmission delay is predicted based on the historical delay mean and the current jitter variance in the link state sampling data. The scheduling weight factor increases as the weighted average of the delay margin decreases; the forward error correction coding parameters are dynamically configured according to the error correction redundancy, and the data units in the coding optimization stream are divided into check blocks. The size of the check block is determined according to the content complexity of the data unit. The content complexity is calculated by combining the texture entropy value and the average motion vector amplitude of the data unit. The higher the content complexity, the smaller the check block. Then, check data is generated according to the error correction redundancy and embedded in the data unit; the dynamically scaled bitrate data, the priority-adjusted scheduling queue, and the data units with embedded check data are reorganized into transmission resources to generate a calibration transmission stream. Each data unit in the calibration transmission stream carries a transmission strategy calibration mark to instruct the receiving end to apply the corresponding decoding compensation strategy. The calibration transport stream is decoded and time-aligned to generate a visual call output sequence that is time-synchronized with the original video stream units, and then pushed to the receiving end presentation device.
2. The 5G-based visual call information processing method according to claim 1, characterized in that, The step of performing spatiotemporal correlation analysis on the data in the network state time series database to identify the periodic patterns and sudden disturbance characteristics of network state fluctuations and generate a network state feature map includes: The data in the network state time series database is decomposed into a time series to separate the trend term, periodic term and random disturbance term. The trend term reflects the long-term change direction of the network state, the periodic term reflects the periodic fluctuation characteristics of the network state, and the random disturbance term represents the state change caused by sudden factors. Spectral analysis is performed on the periodic term to determine the main periodic components and corresponding amplitude characteristics of network state fluctuations. The time-domain signal is converted into a frequency-domain signal by Fourier transform, and the peak frequency in the spectrum is identified as the main periodic component. Statistical distribution analysis is performed on the random disturbance term to calculate the probability density function and cumulative distribution function of the disturbance amplitude, and the threshold condition for the occurrence of the disturbance is determined. When the disturbance amplitude exceeds the threshold, it is determined to be a sudden disturbance event. The spatiotemporal dimensions of the network state feature map are constructed by combining the trend terms, main periodic components and sudden disturbance events. The time dimension is marked with the network state change curve in units of sampling period, and the spatial dimension is plotted as a scatter plot of state distribution with network resource occupancy and signal quality as coordinate axes. In the network state feature map, periodic regular regions and sudden disturbance regions are marked. The periodic regular regions are defined by the confidence interval of the periodic components, and the sudden disturbance regions are marked by the occurrence time and impact range of the disturbance event, thus obtaining a complete network state feature map.
3. The 5G-based visual call information processing method according to claim 2, characterized in that, The step of constructing the state transition probability matrix based on the network state feature map includes: The network states in the network state feature map are divided into multiple discrete state levels. The continuous network state data is classified by a clustering algorithm to determine the boundary of the state level division. The number of transitions between different state levels is counted, and the probability value of transitioning from one state level to another is calculated. The probability value is the ratio of the number of transitions to the total number of transitions, and the initial state transition probability matrix is obtained. The duration distribution of each state level is analyzed, and the average duration and variance of the state level are calculated using survival analysis. A duration distribution model is then established, which is used to describe the stability characteristics of the state level in the time dimension. Combining the initial state transition probability matrix and the duration distribution model, a state transition probability matrix containing two parameters, transition probability and duration, is constructed. Each element in the matrix contains the transition probability from the current state to the target state and the average duration in the current state. The state transition probability matrix is normalized so that the sum of the probabilities of each row element is 1. The convergence of the matrix is verified by a Markov chain model. The construction is completed when the matrix satisfies the stationary distribution condition.
4. The 5G-based visual call information processing method according to claim 1, characterized in that, The step of performing content saliency analysis on the real-time video frame sequence to identify regions of interest and non-regions of interest in the video frames includes: Color features are extracted from each video frame in the real-time video frame sequence. The video frame is converted from the RGB color space to the LAB color space, and the luminance component and chrominance component are separated. The standard deviation of the chrominance component is calculated as the color contrast feature. Texture features are extracted from video frames, and the energy, entropy, contrast, and correlation parameters of the texture are calculated using the gray-level co-occurrence matrix. The texture features are used to characterize the distribution of detailed regions in the video frame. Analyze the motion characteristics of video frames, calculate the pixel displacement vector between adjacent video frames using an optical flow estimation algorithm, and identify regions whose motion vector magnitude exceeds a preset motion threshold as potential regions of interest. Color contrast features, texture features, and motion features are input into a visual attention model and fused by weighted summation. The weight coefficient of motion features within a potential region of interest is set higher than that outside the region to generate a saliency map of the video frame. The value of each pixel in the saliency map represents the visual saliency of that pixel. The saliency map is binarized, and pixel regions with saliency values higher than a threshold are marked as regions of interest, while regions with saliency values lower than the threshold are marked as regions of non-interest.
5. The 5G-based visual call information processing method according to claim 1, characterized in that, The step of dynamically scaling the transmission bitrate of each data unit in the encoded optimized stream according to the bitrate adjustment coefficient, and determining the bitrate adjustment range by multiplying the bitrate adjustment coefficient by the content importance score of the data unit, includes: Region mapping is performed on data units in the coded optimization stream to determine the spatial coordinate range of each data unit in the video frame, extract the saliency value within the coordinate range, and obtain the coding priority parameter corresponding to the data unit; The significance value and the coding priority parameter are weighted and fused. During the fusion process, the weight coefficient of the significance value is higher than that of the coding priority parameter. The sum of the two weight coefficients is 1, and the content importance score of the data unit is obtained. The value range of the content importance score is consistent with the value range of the coding priority parameter. Calculate the bitrate adjustment range, which is equal to the product of the bitrate adjustment coefficient and the content importance score; The current transmission bitrate of the data unit is dynamically scaled based on the bitrate adjustment range. When the bitrate adjustment coefficient is greater than the preset adjustment threshold, the current transmission bitrate is reduced by the bitrate adjustment range. The reduced bitrate value is not lower than the preset minimum bitrate threshold. The preset minimum bitrate threshold is determined based on the basic resolution requirements of the video frame. The scaled bitrate values are smoothly transitioned by setting an upper limit for the bitrate change gradient between adjacent data units. The upper limit for the bitrate change gradient is determined by the ratio of the average bitrate of the encoded optimized stream to the frame rate.
6. The 5G-based visual call information processing method according to claim 1, characterized in that, The step of decoding and aligning the calibration transport stream to generate a visual call output sequence synchronized with the original video stream units includes: The calibration transport stream is received, the transmission strategy calibration mark and compressed data segment of each data unit are extracted, the strategy index is determined by the mark field, and the corresponding decoding compensation strategy is called. The decoding compensation strategy includes error hiding algorithm, motion compensation mode and texture restoration coefficient. Entropy decoding is performed on the compressed data segment. The context model is switched according to the encoding type of the data unit to convert the binary code stream into a quantization coefficient matrix. The encoding type is determined by the start code identifier of the compressed data segment. The quantization coefficient matrix is subjected to inverse quantization and inverse transform processing. The inverse quantization step size in the decoding compensation strategy is adopted. The inverse quantization step size is corrected by the quantization adjustment coefficient in the transmission strategy calibration mark. The transform kernel function is selected according to the intra-frame coding mode or inter-frame coding mode for inverse transform. The motion compensation mode predicts and compensates the pixel data after inverse transformation. In the inter-frame coding mode, the forward or backward reference frames in the reference frame buffer are called and pixel block matching is performed based on the motion vector. During the matching process, a sub-pixel interpolation algorithm is used to improve the accuracy. Loop filtering is performed on the compensated pixel data. The filtering intensity is dynamically adjusted according to the texture restoration coefficient. The higher the texture restoration coefficient, the larger the filter kernel size, thus eliminating block artifacts and ringing effects. The timestamps of the extracted video frames are compared with the system reference clock to calculate the deviation value. When the deviation value is positive, a transition compensation frame is generated by fusing the motion vectors of adjacent frames. When the deviation value is negative, non-critical frames are filtered out and discarded. Non-critical frames are determined by frame type identifier and content importance score. The processed video frames are stitched together in timestamp order to generate a visual call output sequence. Time synchronization accuracy is ensured by using a time sequence check code, which is calculated by the difference between the timestamps of adjacent frames.
7. A computer system, characterized in that, include: processor; And a memory, wherein the memory stores computer-readable code, which, when executed by the processor, causes the processor to perform the 5G-based visual call information processing method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Network adaptive video conference transmission optimization method and system
CN120358347A
Self-adaptive video coding method, device, equipment and medium
CN120390089A