A multi-person oriented video stream parallel processing method and system
By performing quantitative feature analysis and K-means clustering on multiple video streams and dynamically evaluating their importance weights, the problem of uneven resource allocation in parallel processing of multiple video streams is solved, thereby improving the transmission quality and bandwidth utilization of the video streams.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 光谷技术有限公司
- Filing Date
- 2026-03-26
- Publication Date
- 2026-07-07
AI Technical Summary
Existing technologies cannot adapt to real-time changes in video content during parallel processing of multiple video streams, resulting in uneven resource allocation, easy stuttering of critical dynamic video streams, and low overall bandwidth utilization.
By collecting the quantitative characteristics of video streams, K-means clustering analysis is performed to dynamically evaluate the importance weight of groups, monitor transmission delay and stuttering frequency, adjust bitrate allocation, and perform resource reallocation and optimized scheduling.
It achieves adaptive scheduling based on differences in video content, improves resource utilization efficiency, reduces video stream stuttering, and enhances overall transmission quality.
Smart Images

Figure CN122349017A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video communication and real-time transmission technology, and in particular to a method and system for parallel processing of video streams for multiple users. Background Technology
[0002] Currently, real-time video interaction for multiple users has become a core application in fields such as remote collaboration, online education, video conferencing, and interactive live streaming. Its transmission quality and smoothness directly affect user experience and collaboration efficiency. With the increase in the number of access parties and the increasing complexity of application scenarios, efficient parallel encoding and transmission processing of multiple video streams has become a key technical aspect to ensure service quality.
[0003] One existing technology primarily employs a multi-stream video processing scheme based on fixed priority rules or average resource allocation. This scheme typically allocates nearly equal video encoding bitrates and computational resources to each video stream based on preset static roles or simple round-robin scheduling algorithms, ignoring the real-time dynamic changes and complexity differences of the content in each video stream. This method is essentially a static processing mechanism relying on prior settings or simple balancing strategies, and its resource scheduling logic is disconnected from the actual real-time requirements of the video content.
[0004] This processing method, which relies on fixed rules or average allocation, has inherent technical flaws. In multi-channel parallel transmission scenarios, the motion intensity, image detail, and content importance of each video stream dynamically change over time, exhibiting significant differences and imbalances. Static allocation strategies cannot adapt to these real-time changes, leading to insufficient bitrate, stuttering, and image quality degradation when some highly dynamic and important video streams require more resources. Simultaneously, too many resources are allocated to low-dynamic and non-critical video streams, resulting in wasted bandwidth and computing resources. The root cause lies in the method's lack of accurate perception of the real-time content complexity of each video stream, evaluation of dynamic importance weights, and the ability to differentiate and adaptively schedule resource requirements.
[0005] Therefore, the core technical problem faced by existing technologies lies in how to overcome the limitations of static and equal resource allocation to meet the real-time and differentiated requirements of parallel processing of multiple video streams. This can be achieved by analyzing the complexity of video content in real time, dynamically assessing the importance of transmission, and performing adaptive bitrate and priority scheduling based on this, thereby realizing efficient utilization of network resources and optimal balance of overall transmission quality of multiple video streams. Summary of the Invention
[0006] This invention provides a method and system for parallel processing of video streams for multiple users, in order to solve the technical problems of existing technologies that rely on fixed rules or average resource allocation, resulting in the inability to adapt to real-time changes in video content, uneven resource allocation, easy stuttering of key dynamic video streams, and low overall bandwidth utilization in scenarios of concurrent transmission of multiple video streams.
[0007] Firstly, in order to solve the above-mentioned technical problems, the present invention provides a method for parallel processing of video streams for multiple users, comprising: Acquire multiple video streams, extract quantitative features of content changes and motion intensity from each video stream, and generate complexity vectors for each video stream. Based on the complexity vector, K-means clustering analysis is performed on all video streams to group video streams with similar complexity into the same group, thus obtaining a set of grouping categories for the video streams; For each group in the grouping category set, obtain the average complexity and total system bandwidth, and analyze the importance weight of each group based on the average complexity and total system bandwidth; Identify high-priority groups whose importance weights exceed a preset threshold, and adjust the corresponding transmission bitrate according to the real-time network status and coding complexity of the high-priority groups to obtain a bitrate allocation scheme. Monitor the transmission delay and stuttering frequency of key video streams under the bitrate allocation scheme to determine whether there is any waste of resources; If it is determined that there is a waste of resources, then a secondary classification is performed on the high priority group based on the dynamic characteristics of the video stream content to filter out the dynamic content streams that truly require high bitrate and generate a high priority list. If it is determined that there is no waste of resources, then the current bitrate allocation scheme is maintained. Based on the high priority list, bandwidth resources are reallocated to form a transmission quality configuration. Based on the transmission quality configuration, dynamic weights of importance are calculated for all video streams, and transmission requests are arranged in descending order of the dynamic weights to determine the parallel transmission order.
[0008] Secondly, the present invention provides a video stream parallel processing system for multiple users, comprising: The feature extraction module is used to acquire multiple video streams, extract quantitative features of each video stream reflecting content changes and motion intensity, and generate complexity vectors for each video stream. The clustering and grouping module is used to perform K-means clustering analysis on all video streams based on the complexity vector, grouping video streams with similar complexity into the same group to obtain a set of grouping categories for the video streams; The weight calculation module is used to obtain the average complexity and total system bandwidth for each group in the grouping category set, and to analyze the importance weight of each group based on the average complexity and total system bandwidth. The bitrate optimization module is used to identify high-priority groups whose importance weight exceeds a preset threshold, and adjust the corresponding transmission bitrate according to the real-time network status and encoding complexity of the high-priority groups to obtain a bitrate allocation scheme. The resource monitoring module is used to monitor the transmission delay and stuttering frequency of key video streams under the bitrate allocation scheme, and to determine whether there is any waste of resources. The list refinement module is used to perform secondary classification based on the dynamic characteristics of the video stream content from the high priority group if it is determined that there is a waste of resources, to filter out the dynamic content streams that truly need high bitrate, and generate a high priority list. If it is determined that there is no waste of resources, the current bitrate allocation scheme is maintained. The configuration equalization module is used to reallocate bandwidth resources according to the high priority list to form a transmission quality configuration; The scheduling and sorting module is used to calculate the dynamic weight of importance for all video streams according to the transmission quality configuration, and arrange the transmission requests in descending order of the dynamic weight to determine the parallel transmission order.
[0009] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention extracts the complexity vectors of multiple video streams in real time and dynamically calculates the importance weights of groups based on K-means clustering and system bandwidth, thus constructing a content-aware adaptive grouping and priority scheduling mechanism. The technical derivation of this method lies in the fact that existing fixed or equal allocation strategies cannot respond to the differentiated needs brought about by real-time changes in video content. This invention quantifies and clusters the spatiotemporal features of the video and dynamically evaluates the importance of each group in combination with bandwidth constraints, thereby enabling precise allocation of resources to high-complexity, high-demand groups, effectively solving the problems of resource misallocation and critical stream quality degradation caused by ignoring content differences.
[0010] (2) This invention identifies resource waste by monitoring transmission quality and uses support vector machines to refine high-priority groups. Combined with bitrate lower limit adjustment and resource reallocation, it constructs a refined resource balancing model based on feedback optimization. The technical derivation of this method lies in the fact that initial resource tilting may lead to over-allocation. This invention identifies oversaturated streams by introducing a quality saturation baseline and uses machine learning methods to separate out truly high-dynamic content. Subsequently, it dynamically reclaims and reallocates redundant bandwidth, achieving a sublimation from "extensive tilting" to "precise protection," thereby significantly improving the overall utilization efficiency of limited bandwidth resources and avoiding resource waste.
[0011] (3) This invention achieves global optimization of the system-level parallel transmission order by calculating dynamic weights of fused transmission constraints for all video streams before final scheduling and performing topology analysis and decoupling sorting based on the dependency graph. The technical derivation of this method lies in the fact that simple priority queues cannot handle inter-stream dependencies and scheduling blockages. By comprehensively considering deadlines, buffer states, and logical dependency paths, this invention transforms the scheduling problem into a constrained optimization sorting, which can effectively reduce the overall scheduling difficulty of the system, thereby ensuring the orderly and efficient processing of transmission requests in a multi-concurrent environment and improving the system's throughput and real-time performance. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the parallel video stream processing method for multiple users provided in the first embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of the parallel video stream processing method for multiple users provided in the second embodiment of the present invention. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] Reference Figure 1 The first embodiment of the present invention provides a method for parallel processing of video streams for multiple users, including the following steps: S11: Acquire multiple video streams, extract quantitative features of content changes and motion intensity from each video stream, and generate complexity vectors for each video stream. S12, Based on the complexity vector, perform K-means clustering analysis on all video streams, group video streams with similar complexity into the same group, and obtain a set of group categories for the video streams; S13, for each group in the grouping category set, obtain the average complexity and total system bandwidth, and analyze the importance weight of each group based on the average complexity and total system bandwidth; S14, identify high-priority groups whose importance weight exceeds a preset threshold, and adjust the corresponding transmission bitrate according to the real-time network status and encoding complexity of the high-priority groups to obtain a bitrate allocation scheme. S15, monitor the transmission delay and stuttering frequency of key video streams under the bitrate allocation scheme, and determine whether there is any waste of resources; S16. If it is determined that there is a waste of resources, then from the high priority group, a secondary classification is performed based on the dynamic characteristics of the video stream content to filter out the dynamic content streams that truly need high bitrate and generate a high priority list. If it is determined that there is no waste of resources, then the current bitrate allocation scheme is maintained. S17, According to the high priority list, bandwidth resources are reallocated to form a transmission quality configuration; S18, based on the transmission quality configuration, calculate the dynamic weight of importance for all video streams, and arrange the transmission requests in descending order of the dynamic weight to determine the parallel transmission order.
[0015] In step S11, multiple video streams are acquired, and quantitative features reflecting content changes and motion intensity are extracted from each video stream to generate a complexity vector for each video stream, including: Acquire multiple video streams and calculate the grayscale difference between adjacent sampled frames, and generate inter-frame difference data based on the grayscale difference; Macroblock matching is performed on the inter-frame difference data to extract displacement vectors, and the sum of the magnitudes of the displacement vectors is calculated to generate motion intensity data. Video frames are processed using edge detection algorithms to extract image edge gradient information; The motion intensity data and the image edge gradient information are input together into a preset feature mapping model to extract quantitative feature values that reflect spatiotemporal complexity. The deviation is calculated based on the statistical distribution characteristics of the quantized feature values, and the deviation is normalized to obtain the complexity vector of each video stream.
[0016] First, multiple video streams are accessed, and the grayscale difference between adjacent sampled frames is calculated to generate inter-frame difference data. These multiple video streams originate from real-time videos uploaded simultaneously by multiple participants in scenarios such as video conferencing, online education, or live interactive sessions. For each video stream, the system acquires a continuous sequence of images at a fixed frame rate. To quantify the content changes between adjacent frames, each frame is converted to a grayscale image, and the absolute difference in grayscale values between temporally adjacent frames is calculated pixel by pixel. For a frame with a resolution of 1920x1080, a data matrix containing 2,073,600 differences is obtained. By calculating the arithmetic mean of all elements in the entire difference matrix, a scalar value representing the intensity of inter-frame change at that moment is obtained. These scalar values are continuously calculated and recorded, thus forming an inter-frame difference data sequence describing the intensity of the video content change over time.
[0017] Subsequently, macroblock matching is performed on the inter-frame difference data to extract displacement vectors, and motion intensity data is generated accordingly. To more finely depict motion within the frame, each video frame is divided into multiple fixed-size macroblocks; for example, 16x16 pixel macroblocks are used. For two consecutive frames, in the second frame, for each macroblock of the first frame, a matching block with the smallest grayscale difference is searched within a preset search window. The matching process uses absolute error as a similarity criterion. The positional offset between the matching block and the original macroblock constitutes the displacement vector of that macroblock. By summing the magnitudes of the displacement vectors of all macroblocks within a frame, a motion intensity value characterizing the overall motion intensity of that frame can be obtained.
[0018] Next, video frames are processed using an edge detection algorithm to extract image edge gradient information. Motion intensity data and image edge gradient information are input into a feature mapping model to extract quantized feature values reflecting spatiotemporal complexity. Motion intensity data reflects dynamic changes in the temporal dimension, while image edge gradient information reflects structural detail complexity in the spatial dimension. Edge gradient information can be obtained by applying the Sobel edge detection operator to video frames and calculating the edge pixel density of the entire frame. The feature mapping model is a pre-defined nonlinear function. This function normalizes the motion intensity values... With the normalized edge gradient density value The fusion outputs a comprehensive quantized feature value. Its expression is: .in and Before input, each input element has been normalized using its historical maximum and minimum values, mapping it to the [0,1] interval to eliminate the influence of dimensions. Weighting coefficients and To determine the optimal encoding method through historical data regression analysis, a large number of video clips covering different motion and texture complexities were collected. Their overall complexity was labeled according to subsequent encoding bitrate requirements, and the optimal method was fitted using linear regression. and The value that makes the function output value This value has the highest correlation with the complexity label. It takes into account both the intensity of motion and the richness of texture details in the image, and can more comprehensively measure the difficulty of encoding and transmitting video content.
[0019] Then, the deviation is calculated based on the statistical distribution characteristics of the quantized feature values, and normalization is performed to obtain the final complexity vector. Since different video sources have different content characteristic baselines, directly using the original quantized feature values is not conducive to fair comparison across video streams. Therefore, a short-term statistical benchmark needs to be established for each video stream. The system continuously records all quantized feature values of each video stream within a recent time window and calculates the mean of the feature values within that time window. and For the latest quantized eigenvalues Calculate its deviation from this local statistical distribution. Calculated using standard scores. This deviation This characterizes the fluctuation of the current content complexity relative to its recent normal. Positive values indicate that the current complexity is higher than the recent average, while negative values indicate that it is lower than the average. Finally, the deviation... The data is normalized using the Sigmoid function to map it to the (0,1) interval. This normalized value becomes one dimension element of the complexity vector. By continuously executing the above process, the system generates such a complexity scalar value for each video stream at fixed time intervals, and arranges them in chronological order, ultimately forming the complexity vector of that video stream over a period of time. This vector abandons absolute complexity and instead uses relative fluctuation to characterize the dynamic changes in the video stream's resource requirements, providing stable and comparable input features for subsequent clustering and grouping.
[0020] In step S12, based on the complexity vector, K-means clustering analysis is performed on all video streams to group video streams with similar complexity into the same group, resulting in a set of video stream grouping categories, including: Select initial cluster centroids from the complexity vector; Calculate the Euclidean distance between each complexity vector and the initial cluster centroid, and assign each complexity vector to a temporary group with the minimum Euclidean distance; Update the cluster centroids based on the mean of the complexity vectors within each temporary group, and iterate the assignment and update operations until the centroid changes converge. A group index table is established based on the converged clustering relationships to obtain the set of group categories for the video stream.
[0021] First, the complexity vectors of all video streams are obtained, and initial cluster centroids are selected. The complexity vectors, generated in step S11, are one-dimensional time-series vectors representing the dynamic changes in the content of each video stream. Before performing cluster analysis, the number of clusters needs to be determined. This value The initial cluster centroids are not fixed and can be determined based on prior knowledge of the business scenario or by analyzing the silhouette coefficients of historical data using the elbow rule. The selection of the initial cluster centroids affects the algorithm's convergence speed and the final result. This implementation uses an improved K-means++ algorithm for initialization to reduce the risk of getting trapped in local optima.
[0022] Subsequently, the Euclidean distance between each complexity vector and all initial centroids is calculated, and it is assigned to a temporary group according to the nearest neighbor principle. Next, based on the members within the current temporary group, the cluster centroid of each group is recalculated, and the assignment and update operations are iteratively performed until the convergence condition is met. For each temporary group, the arithmetic mean of all complexity vectors within it is calculated on each feature dimension, and the new vector formed by this mean is used as the updated centroid of that group. In all... After all centroids have been updated, the algorithm enters the next iteration, recalculates the Euclidean distance of all complexity vectors to the new centroid set, and reallocates them to temporary groups accordingly.
[0023] Finally, based on the final clustering relationships determined after convergence, a structured group index table is established, thereby obtaining the set of group categories for the video streams. The system records the unique group number to which each video stream ultimately belongs. The group index table is a key-value pair structure or database table, where the key is the video stream identifier and the value is the group number to which the video stream belongs.
[0024] In step S13, for each group in the grouping category set, the average complexity and total system bandwidth are obtained, and the importance weight of each group is analyzed based on the average complexity and total system bandwidth, including: The average motion vector magnitude and average edge pixel density of each video stream group are obtained as the average complexity. The average motion vector magnitude and the average edge pixel density are weighted and fused to obtain the overall complexity. Calculate the ratio of the overall complexity to the current total system bandwidth capacity to obtain the relative bandwidth requirement percentage; If the relative bandwidth demand ratio exceeds the preset congestion warning threshold, the relative bandwidth demand ratio is non-linearly amplified to determine the initial importance weight value. Based on the preliminary importance weight values, the video streams are sorted from high to low to determine their priority order.
[0025] First, the average motion vector magnitude and average edge pixel density of each video stream group are obtained as core data characterizing the overall content complexity of the group. The group category set is generated in step S12, where each group contains several video streams with similar complexity. For each group, the system iterates through all video streams within it, calculating the motion intensity data and edge gradient data within the most recent statistical window. The average motion vector magnitude is obtained by calculating the average of the motion intensity data of each frame of all video streams within the group within the statistical window; this value reflects the overall intensity of dynamic changes in the video within the group. The average edge pixel density is obtained by calculating the average ratio of the number of edge pixels to the total number of pixels in each frame of all video streams within the group after edge detection within the statistical window; this value reflects the overall richness of detail and texture in the video within the group. For example, in a video conferencing scenario, a group with multiple participants speaking and gesturing may have an average motion vector amplitude of 80 pixels per second and an average edge pixel density of 0.15; while a group with multiple participants with static avatars or sharing static documents may have an average motion vector amplitude of only 5 pixels per second and an average edge pixel density of 0.08.
[0026] Subsequently, the obtained average motion vector magnitude and average edge pixel density are weighted and fused to calculate the comprehensive content complexity, which characterizes the overall encoding and transmission difficulty of the group. Specifically, the average motion vector magnitude and average edge pixel density are first subjected to min-max normalization to scale them to the same numerical range (e.g., [0,1]) to eliminate the influence of differences in the original units and numerical ranges. Then, the comprehensive content complexity is calculated using a weighted summation method. The weighting coefficients are pre-set according to the relative importance of motion and detail to video quality and bandwidth requirements in different application scenarios. For example, in real-time game scenarios that emphasize smoothness, motion has a higher weight, and the weighting coefficient can be set to 0.7:0.3; in medical imaging scenarios that emphasize clarity, detail has a higher weight, and the weighting coefficient can be set to 0.3:0.7.
[0027] Next, the ratio of the overall content complexity of each group to the total available bandwidth of the current system is calculated to obtain the relative bandwidth demand ratio of each group. This ratio quantifies the relative urgency of the bandwidth required to meet the expected quality of the video streams in that group under the current network resource conditions. The total system bandwidth capacity is the available bandwidth obtained through real-time monitoring. Then, it is determined whether the relative bandwidth demand ratio of each group exceeds a preset congestion warning threshold. This threshold is set to identify the risk threshold that may cause network congestion or quality degradation. Its determination method can be based on historical service quality data statistics. For example, by analyzing historical logs, the inflection point where the probability of a sharp increase in stuttering or latency in the corresponding group's video streams increases dramatically when the relative bandwidth demand ratio exceeds a certain value can be found, and this value is set as the threshold. If the relative bandwidth demand ratio of a group is greater than this threshold, it indicates that the bandwidth demand of that group is already urgent under the current network conditions and requires higher attention. At this time, the relative bandwidth demand ratio is non-linearly amplified to highlight its importance. One possible amplification function is... ,in The parameter is greater than 0, and its value is determined by analyzing the correlation between the proportion of bandwidth demand in historical data and the subsequent actual congestion or quality degradation, to ensure that a significant weighting effect is generated when the demand proportion exceeds the warning threshold. This is the preliminary importance weight value obtained from the calculation. This represents the percentage of relative bandwidth demand. This is the congestion warning threshold. This non-linear amplification ensures that the high-demand group receives an exponentially increasing weighting in resource competition. Through this step, the weight of the high-dynamic group may be amplified from 0.075 to 0.15, while the weight of the low-dynamic group remains unchanged at 0.025.
[0028] Finally, based on the calculated preliminary importance weight values of all groups, they are sorted in descending order. The sorting result directly determines the priority order of each video stream group.
[0029] In step S14, high-priority groups whose importance weights exceed a preset threshold are identified. Based on the real-time network status and coding complexity of these high-priority groups, the corresponding transmission bitrate is adjusted to obtain a bitrate allocation scheme, including: If the importance weight of a certain video stream group exceeds a preset importance threshold, it is determined to be a high-priority group; Collect real-time buffer occupancy and packet loss rate data of the high-priority group and calculate the network congestion status value; Obtain the average motion intensity and texture complexity of the high-priority group, and calculate the encoding complexity based on the average motion intensity and texture complexity; The network congestion status value and the coding complexity of the high priority group are input into a preset nonlinear mapping function to obtain the incremental allocation coefficient; The incremental allocation coefficient is superimposed on the preset basic allocation ratio to generate the corrected allocation ratio; The target transmission bitrate is calculated based on the corrected allocation ratio, and the quantization parameters are derived based on the target transmission bitrate to obtain the bitrate allocation scheme.
[0030] First, based on the importance weights of each video stream group calculated in step S13, high-priority groups requiring priority resource allocation are identified. The importance threshold is determined based on statistical analysis of historical operational data. The distribution of weight values for each group during stable system operation is collected, and a higher quantile (e.g., the 95th quantile) is selected as the threshold. This method ensures that, in most cases, only groups with significantly higher demand than normal levels are identified as high-priority. This avoids both setting the threshold too low, leading to misjudgment of too many groups, and setting the threshold too high, causing omissions of critical needs. The system iterates through all groups; if the importance weight of a group is greater than the threshold, that group is designated as a high-priority group.
[0031] Subsequently, for the identified high-priority groups, real-time network transmission status data, including buffer occupancy and packet loss rate, is collected, and a comprehensive network congestion status value is calculated. Specifically, for each video stream within the group, its receiver buffer occupancy percentage and recent packet loss rate (Loss) are obtained through feedback information from the Transmission Control Protocol (TCP) or Real-Time Transport Protocol (RTP). The average of these two metrics for all streams within the group is then calculated to obtain the average buffer occupancy rate for that group. and average packet loss rate Network congestion status values can be calculated using a weighted summation. The weights are determined based on historical data analysis. Specifically, the correlation coefficients between buffer occupancy rate and packet loss rate and the eventual occurrence of perceptible stuttering or quality degradation are statistically analyzed across a large number of historical transmission segments. These correlation coefficients are then normalized and used as weights. For example, if the analysis shows that buffer occupancy rate has a stronger predictive correlation with quality degradation, a higher weight, such as 0.6, is assigned to it. Next, the calculated network congestion status value and the coding complexity of the high-priority group are input into a non-linear mapping function to calculate the bandwidth increment allocation coefficient for that group. The encoding complexity is obtained by calculating the weighted sum of the average motion intensity and average texture complexity of all video streams within the high-priority group, and its specific calculation method is consistent with the comprehensive content complexity value described in step S13. For example, the nonlinear mapping function is in the form of:
[0032] in, This represents the baseline adjustment range under no additional pressure, and is usually set to a small empirical value. and These represent the "normal" levels of network conditions and content complexity, respectively, and can be obtained by statistically analyzing the median of historical data. It is the hyperbolic tangent function; and For sensitivity coefficient control and The strength of the impact of deviations from the baseline on the increment is determined through a supervised parameter tuning process: a historical dataset is constructed, where each sample contains a set of... Input values and, given this set of inputs, different candidate values. The overall system quality evaluation index corresponds to the bitrate allocation scheme calculated from the parameters. Optimization methods such as grid search or gradient descent are used to find the scheme that optimizes the overall quality evaluation index on this historical dataset. Parameter combinations.
[0033] For example, suppose the current network congestion status value of a high-priority group is... =0.8, baseline value =0.5, coding complexity =0.7, and set =0.1, =2, =1. Calculate the linear term: .through After mapping, substituting into the formula yields... .
[0034] Then, the calculated incremental allocation coefficients This is superimposed on the original base allocation ratio for the high-priority group. A revised bandwidth allocation ratio is generated. The new allocation ratio equals the original ratio plus the increment factor. The product of [the product of the two groups]. Meanwhile, to ensure that the total bandwidth allocation does not exceed 100%, the allocation ratio for other non-high-priority groups needs to be reduced accordingly.
[0035] Finally, the target total bandwidth for the high-priority group is calculated by multiplying the corrected bandwidth allocation ratio by the current total available bandwidth of the system. Subsequently, the target total bandwidth is allocated to each video stream within the group, based on the original importance weights calculated in step S13. For example, if there are three video streams in the group with a normalized weight ratio of 0.5:0.3:0.2, the target total bandwidth is also allocated to these three streams in this ratio. After obtaining the target transmission bitrate for each video stream, the video encoder adjusts encoding control parameters such as quantization parameters using a bitrate control algorithm to achieve that bitrate. This complete scheme, from allocation ratio to encoding parameter adjustment, constitutes an optimized bitrate allocation scheme tailored to the current network conditions and content requirements.
[0036] In step S15, the transmission delay and stuttering frequency of key video streams under the bitrate allocation scheme are monitored to determine whether there is resource waste, including: Obtain actual transmission delay data and stuttering frequency data for key video streams; The actual transmission delay data and the stuttering frequency data are mapped to a preset quality saturation baseline to identify oversaturated time segments where the transmission quality is higher than the baseline. Calculate the resource overflow magnitude within the oversaturation time segment. If the resource overflow magnitude is greater than a preset redundancy tolerance value, it is determined that there is resource waste; otherwise, it is determined that there is no resource waste.
[0037] First, the actual transmission performance data of the identified critical video streams under the bitrate allocation scheme is obtained, specifically including actual transmission delay data and stuttering frequency data. The critical video streams typically refer to video streams in high-priority groups or important streams designated by business logic. Actual transmission delay data is obtained by timestamping each frame of data at the sending end and calculating the difference between the receiving time and the sending time at the receiving end; this data reflects the end-to-end delay from data transmission to reception. Stuttering frequency data is obtained by monitoring the continuity of video playback at the receiving end; when the time interval between two frames exceeds a preset smooth playback threshold, it is recorded as a stutter. The smooth playback threshold is typically set to 1.5 to 2.5 times the frame interval corresponding to the video frame rate. The number of stutters occurring per unit time is counted to obtain the stuttering frequency data. The system continuously collects and records this time-series data, providing a basis for subsequent analysis.
[0038] Subsequently, the acquired actual transmission delay data and stuttering frequency data are mapped to a preset quality saturation baseline to identify oversaturated time segments where transmission quality exceeds the necessary level. The quality saturation baseline defines the minimum performance standard required to ensure an acceptable user experience, typically including an upper limit for delay and an upper limit for stuttering frequency. This baseline is set based on statistical analysis of historical service quality data. Specifically, a large amount of transmission performance data from historical sessions where users did not report stuttering or delay complaints is collected. The distributions of delay and stuttering frequency are calculated separately, and the lower quantile of the delay distribution (e.g., the 10th quantile) is selected as the upper limit for delay, and the lower quantile of the stuttering frequency distribution (e.g., the 10th quantile) is selected as the upper limit for stuttering frequency. The mapping process involves dividing the continuous monitoring time axis into fixed-length time segments. For a given time segment, the average actual transmission delay and the average stuttering frequency of all sampling points within it are calculated. If both the average actual transmission delay and the average stuttering frequency are below the upper limit for delay and stuttering frequency respectively, the transmission capacity of that time segment is determined to be higher than the baseline requirement, i.e., it is in an oversaturated state. All time segments that meet this condition are marked as edge-saturated time segments.
[0039] Next, for each identified oversaturated time segment, the magnitude of the resource overflow is calculated. This value quantifies the extent to which the bandwidth allocated to the video stream exceeds its actual demand within that time period. One calculation method is:
[0040] in, The value representing the resource overflow is between 0 and... The value is a dimensionless value between the baseline and the given value. A larger value indicates that the actual mass exceeds the baseline more significantly, and the greater the potential for resource waste. This represents the average actual transmission delay within the oversaturated time segment. This is the preset upper limit baseline value for delay; This represents the average frequency of stuttering within the oversaturated time segment. This is the preset upper limit for the frequency of stuttering; and These are the weighting coefficients for latency and stuttering frequency in the overflow assessment, respectively, satisfying... The weighting coefficients are determined by analyzing the regression coefficients of latency and stuttering on users' subjective quality ratings in historical data. Specifically, samples containing different levels of latency and stuttering and their corresponding subjective ratings are collected, and multiple linear regression analysis is performed. The normalized regression coefficients are then used as the weighting coefficients. and The value of .
[0041] Then, the calculated resource overflow magnitude value With the preset redundancy tolerance value The redundancy tolerance value is compared. This defines the acceptable performance margin boundaries reserved for the system to ensure stability. These boundaries are based on the statistical distribution of long-term monitoring data, collecting all calculated resource overflow values during long-term system operation and analyzing their distribution. This is set to a relatively high quantile of the distribution, such as the 90th quantile. This means that the system, by default, allows a certain proportion of oversaturated states to not be considered as resource waste that needs to be dealt with immediately. This provides a buffer for normal fluctuations in network conditions and instantaneous bitrate fluctuations in the encoder, avoiding frequent resource reallocation due to excessive sensitivity, which would affect system stability.
[0042] Finally, if the resource overflow magnitude exceeds a preset redundancy tolerance value, the decision logic determines that the critical video stream exhibits resource waste within the current monitoring context. This means that the allocated bandwidth resources significantly exceed the level required to maintain baseline quality, and there is room to reclaim some resources to allocate to other video streams that require more bandwidth. The decision result will trigger subsequent optimization steps. If this condition is not met, it is determined that there is no significant resource waste, and the current resource allocation is maintained.
[0043] In step S16, if resource waste is determined to exist, a secondary classification is performed from the high-priority group based on the dynamic characteristics of the video stream content to filter out the dynamic content streams that truly require high bitrates, generating a high-priority list. If resource waste is determined not to exist, the current bitrate allocation scheme is maintained, including: If it is determined that there is a waste of resources, then multi-dimensional spatiotemporal feature data of each video stream is extracted from the high-priority group; Based on the multidimensional spatiotemporal feature data, a feature matrix to be classified is constructed. The feature matrix to be classified is input into a preset support vector machine model to calculate the geometric interval value. Based on the geometric interval value, vectors located on the positive side of the classification hyperplane boundary are selected, and the video streams corresponding to the selected vectors are marked as dynamic content streams that truly require high bitrate. Based on all the marked dynamic content streams, the high priority list is generated. If it is determined that there is no waste of resources, the bandwidth allocation ratio and transmission bitrate of each video stream group will remain unchanged, and the current bitrate allocation scheme will be maintained.
[0044] First, branching is performed based on the judgment result of step S15. If resource waste is determined, the system initiates a refined screening process for high-priority groups. The goal of this process is to further distinguish from the groups that have been macroscopically marked as "high priority" those video streams with highly dynamic content and a real and urgent need for high bitrate, and those video streams that, although belonging to high-priority groups, currently have average dynamism and may have excessive resource allocation.
[0045] Subsequently, multidimensional spatiotemporal feature data for each video stream is extracted from the high-priority group. Here, "multidimensional" means that the features simultaneously cover both the spatial and temporal domains. Specifically, for each video stream within the group, the following features are extracted within the most recent analysis period: temporal domain features, such as the average value of inter-frame brightness differences and the variance of motion vector amplitudes; spatial domain features, such as the texture complexity of keyframes and the rate of change in the area proportion of the region of interest in the frame. These features collectively constitute a feature vector describing the dynamic characteristics of the video stream in the spatiotemporal dimensions.
[0046] Next, based on the extracted multidimensional spatiotemporal feature data of all video streams, a feature matrix to be classified is constructed. The rows of this matrix correspond to each video stream in the high-priority group, and the columns correspond to the aforementioned multidimensional spatiotemporal features. This matrix is then input into a pre-trained support vector machine model. This model is a binary classifier, and its training objective is to distinguish between truly high-dynamic content streams and non-high-dynamic content streams. The model's training data comes from historical labeled data, where positive samples are video stream segments that historically benefited from high bitrates due to dramatic dynamics, such as rapid movement or complex scene changes, while negative samples are video stream segments that were assigned high bitrates but whose actual content was dynamically flat, potentially wasting resources.
[0047] During the training phase, the model solves a convex quadratic programming problem using an optimization algorithm. The goal is to find a set of parameters that correctly classifies all training samples and maximizes the classification margin, thus obtaining the optimal hyperplane. In this implementation, a linear kernel function is used, therefore the classification decision function is linear. For a new feature vector to be classified, the support vector machine model determines its classification by calculating the geometric margin to the optimal hyperplane.
[0048] Then, filtering is performed based on the calculated geometric margin values. The system sets a filtering threshold. Vectors located on the positive side of the classification hyperplane boundary, i.e., feature vectors with a geometric margin greater than the threshold, are selected as video streams. These video streams are judged by the model with high confidence to be highly matched to the pattern of "truly high dynamic range content." These video streams are then marked as truly dynamic content streams requiring high bitrates. The threshold setting aims to control the strictness of the filtering; its value can be adjusted on the validation set. A series of candidate thresholds are traversed, and the F1 score between the set of video streams selected by the model and the set of manually labeled true high dynamic range streams is calculated at each threshold. The candidate threshold with the highest F1 score is selected as the final threshold. Afterward, a high-priority list is generated based on all labeled dynamic content streams. For example, the original high-priority group might contain 15 video streams; after filtering by the support vector machine, only 8 might be included in this list.
[0049] Finally, if it is determined that there is no resource waste, the other branch of logic is executed. At this point, it is considered that the current bitrate allocation scheme has achieved a good balance between ensuring quality and resource utilization, and there is no need for radical resource readjustment. Therefore, the system maintains the current bandwidth allocation ratio and transmission bitrate of each video stream group unchanged.
[0050] In step S17, bandwidth resources are reallocated according to the high-priority list to form a transmission quality configuration, including: The remaining video stream objects are filtered according to the high-priority list; Obtain the motion intensity and texture complexity of the remaining video stream object, calculate the encoding complexity based on the motion intensity and texture complexity, and calculate the lower limit of the bitrate based on the encoding complexity; Calculate the difference between the real-time transmission rate of each remaining video stream object and the corresponding lower limit of the bitrate, and sum them up to obtain the total amount of redundant bandwidth resources; The total amount of redundant bandwidth resources is allocated to the video streams in the high-priority list to obtain the enhanced bitrate value of the key streams; The lower bitrate value and the enhanced bitrate value are combined to form a transmission quality configuration.
[0051] First, the remaining video stream objects are filtered based on the high-priority list. This high-priority list, generated in step S16, contains refined identifiers of dynamic content streams that truly require high bitrates. The system excludes members belonging to the high-priority list from the current set of all active video streams; the remaining streams are defined as remaining video stream objects. These objects typically include static background streams, low-dynamic-content streams, or video streams not related to critical business processes. For example, in a video conference, after the speaker and several frequently interacting participants' video streams are included in the high-priority list, the remaining participant video streams, which only display static avatars or shared documents, are classified as remaining video stream objects.
[0052] Subsequently, the lower limit of the bitrate of the remaining video stream is calculated based on the encoding complexity of the remaining video stream object. (Bitrate Lower Limit) This refers to the minimum bitrate required for the video stream while maintaining a basically acceptable quality. The calculation of this value is closely related to the encoding complexity of the video stream. It is determined by quantizing the spatiotemporal entropy of video frames. Specifically, for each video stream, within a fixed time window, the average motion vector magnitude is calculated as the time dimension complexity. And calculate its average image gradient magnitude as the spatial dimension complexity. The final coding complexity The weighted sum of the two. The weighting coefficient This was obtained by regressing the impact of both factors on the encoded bitrate using historical data. A definite... The method is to use a regression model trained on historical data. The model training process involves collecting a large number of video clips covering different content types, encoding them using a specified encoder in a preset constant quality mode, and recording the actual bitrate of each encoded clip as its value. Labels, and simultaneously calculate the encoding complexity of each segment. As a feature, the parameters are analyzed using the least squares method. and By fitting the data, a regression model is obtained.
[0053] Next, the difference between the real-time transmission rate of each remaining video stream object and the corresponding lower limit of the remaining video stream bitrate is calculated, and these differences are summed to obtain the total amount of redundant bandwidth resources. The system monitors the current actual transmission rate of each remaining video stream object in real time, which is usually determined by the bitrate control module at the sending end or obtained by network probing. For each remaining video stream object, the difference between its rate and the lower limit of the bitrate is calculated. If the difference is greater than 0, it indicates that the bandwidth currently occupied by the stream exceeds the minimum value required to maintain basic quality, and there is reclaimable redundancy; if the difference is less than or equal to 0, it indicates that the stream is at or below the minimum guarantee line, and no resource reclamation is performed. The differences of all remaining video stream objects that satisfy the condition of a difference greater than 0 are summed to obtain the total amount of redundant bandwidth resources that the system can currently reclaim.
[0054] Then, the total amount of redundant bandwidth resources is allocated to the video streams in the high-priority list to obtain the enhanced bitrate values for the key streams. The allocation strategy can be weighted allocation, and the weights can be determined by normalizing the geometric interval values output by the support vector machine in step S16. This is because a larger geometric interval value indicates a higher confidence level in the model's judgment that the stream is high-dynamic content, and it should receive more additional bandwidth. The geometric interval values of all allocated video streams are summed, and the normalized allocation ratio for any video stream is its own geometric interval value divided by the total geometric interval value.
[0055] Finally, the remaining video stream bitrate lower limit and the key stream enhanced bitrate values are aggregated to form a transmission quality configuration. For each remaining video stream object, its configured target bitrate is its calculated lower bitrate limit. For each video stream in the high-priority list, its configured target bitrate is its key stream enhanced bitrate. For other video streams that are neither part of the remaining video stream object nor in the high-priority list, their current bitrate allocation can be maintained. By summarizing the identifiers of all video streams and their corresponding new target bitrates, a complete and structured transmission quality configuration is formed.
[0056] In step S18, based on the transmission quality configuration, dynamic weights of importance are calculated for all video streams, and transmission requests are arranged in descending order of the dynamic weights to determine the parallel transmission order, including: Extract the identifiers of each video stream from the transmission quality configuration, and calculate the dynamic importance weights based on the transmission deadline constraints and buffer occupancy rates corresponding to each identifier. The dynamic importance weights of each video stream are mapped to a preset dependency graph, and the logical dependency paths between video streams are determined through topology analysis. The scheduling difficulty coefficient is calculated based on the logical dependency path. If the scheduling difficulty coefficient is higher than the preset difficulty threshold, the logical dependency path is decoupled and a sorted index table is generated. If the scheduling difficulty coefficient is not higher than the preset difficulty threshold, a sorted index table is generated based on the logical dependency path and the dynamic importance weight. The transmission request queues of all video streams are rearranged according to the sorting index table to obtain an ordered queue. The ordered queue is then traversed to determine the final parallel transmission order.
[0057] First, the identifiers of each video stream are extracted from the transmission quality configuration. Based on the transmission deadline constraints and buffer occupancy rates corresponding to each identifier, dynamic importance weights are calculated. The transmission quality configuration is generated in step S17, where the system obtains the transmission deadline constraint and the current receiver buffer occupancy rate for each video stream. The dynamic importance weights are calculated based on these two indicators. To unify parameters with different dimensions and time units into summable scalar values, normalization is required. The remaining deadline is converted into a relative urgency related to the baseline delay, and the buffer occupancy rate is treated as a scaling factor. A specific calculation method is as follows:
[0058] in, The dynamic importance weight is a dimensionless numerical value. Due to transmission deadline constraints, For the current time, The reference time constant is used to dimensionlessly represent the time term. Its value is typically set to a typical, tolerable delay for the system. The introduction of this constant makes the time urgency term a relative ratio independent of specific absolute time values. Its physical meaning lies in mapping the remaining absolute time to a value that is dimensionless. On the basis of relative urgency measurement as a reference scale; It should be a very small positive number to prevent the denominator from being zero; This represents the current receiver buffer occupancy rate, which is dimensionless and usually expressed as a decimal or percentage. and The adjustment weighting coefficients for the time urgency term and the buffer occupancy rate term, respectively, satisfy... The determination of these two coefficients is based on historical data analysis. In historical operational data, the correlation coefficients between time urgency and buffer occupancy rate and the eventual occurrence of playback stuttering or data expiration are calculated. The normalized correlation coefficients are then used as the basis for determining these coefficients. and The value of is determined by the formula. This means that video streams closer to their deadline and with higher buffer occupancy rates have greater dynamic importance weights and should receive higher priority in scheduling.
[0059] Next, the dynamic importance weights of each video stream are mapped to a preset dependency graph, and logical dependency paths between video streams are determined through topological analysis. The dependency graph is a directed graph where nodes represent video streams and edges represent logical dependencies between streams. Dependencies may be generated by application logic; for example, in remote collaboration, shared whiteboard video stream B may depend on the content of speaker-led video stream A for synchronization. In this case, there exists an edge in the graph pointing from node A to node B, indicating that A should be transmitted before or synchronously with B. The system will then calculate the dynamic importance weights. These attributes are assigned to the corresponding nodes in the graph. Topological analysis identifies logical dependency paths by analyzing this dependency structure; that is, a sequence of nodes that starts from a source node and travels along the dependency edge to another node. For example, the analysis might find a path "A->B->C", indicating that C depends on B, and B depends on A.
[0060] Then, the scheduling difficulty coefficient is calculated based on the logical dependency path. If the scheduling difficulty coefficient is higher than a preset difficulty threshold, the logical dependency path is decoupled, and a sorted index table is generated. The scheduling difficulty coefficient is calculated based on critical path analysis. For example, the maximum value of the sum of the dynamic importance weights of nodes on all logical dependency paths in the computation graph is used as the adjustment difficulty coefficient. The preset difficulty threshold is set based on statistical analysis of historical scheduling performance data. The average request processing latency under different scheduling difficulty coefficients during the system's historical operation is collected, and the scheduling difficulty coefficient value corresponding to the inflection point that causes the average latency to begin to increase significantly non-linearly is found and used as the threshold. If the real-time scheduling difficulty coefficient is higher than the preset difficulty threshold, the decoupling operation is triggered. In specific implementation, the system traverses all logical dependency paths. Paths whose length exceeds the preset length threshold and whose sum of dynamic importance weights on the path is high are identified as long dependency critical paths that need to be decoupled. The selection rule for cutting off nodes is based on the dynamic importance weights of each node on the path and its topological position in the path. Nodes with relatively low weights and located in the middle of the path are usually selected as "cut-off points". For example, for a path A→B→C→D, if the dynamic importance weight of node B is significantly lower than that of A and C, then the path is cut off at B, decoupling the original path into two sub-sequences, A→B and C→D, which can be scheduled in parallel or partially overlapped. After the cutoff, the original dependency is partially removed, allowing subsequent nodes to be scheduled earlier under certain conditions. Furthermore, the absolute dependence of downstream nodes on the real-time arrival of data from upstream nodes can be reduced by dynamically adjusting the bitrate, thereby weakening the dependency strength. After decoupling, the system generates a sorted index table based on the new schedulable relationships of the nodes and their dynamic importance weights. This table is an ordered list that clearly defines the order in which the video streams are scheduled and processed.
[0061] Finally, the transmission request queues of all video streams are rearranged according to the sorting index table to obtain an ordered queue. This ordered queue is then traversed to determine the final parallel transmission order. The system rearranges the transmission request queues, which may have been unordered or simply FIFO, based on the generated sorting index table. Video streams that appear earlier in the sorting index table have their transmission requests placed at the front of the ordered queue. Subsequently, the system scheduler traverses the ordered queue, processing the transmission request of each video stream in turn. The traversal order determines the final parallel transmission order of the video stream data packets entering the network transmission channel. This order is an optimized result derived after comprehensively considering multiple factors such as content importance, transmission urgency, buffer status, and inter-stream logical dependencies. It aims to reduce scheduling congestion caused by dependencies, prioritize critical and urgent video streams, and thus achieve better overall transmission efficiency and real-time performance in a parallel environment.
[0062] Reference Figure 2 The first embodiment of the present invention provides a parallel video stream processing system for multiple users, comprising: The feature extraction module is used to acquire multiple video streams, extract quantitative features of each video stream reflecting content changes and motion intensity, and generate complexity vectors for each video stream. The clustering and grouping module is used to perform K-means clustering analysis on all video streams based on the complexity vector, grouping video streams with similar complexity into the same group to obtain a set of grouping categories for the video streams; The weight calculation module is used to obtain the average complexity and total system bandwidth for each group in the grouping category set, and to analyze the importance weight of each group based on the average complexity and total system bandwidth. The bitrate optimization module is used to identify high-priority groups whose importance weight exceeds a preset threshold, and adjust the corresponding transmission bitrate according to the real-time network status and encoding complexity of the high-priority groups to obtain a bitrate allocation scheme. The resource monitoring module is used to monitor the transmission delay and stuttering frequency of key video streams under the bitrate allocation scheme, and to determine whether there is any waste of resources. The list refinement module is used to perform secondary classification based on the dynamic characteristics of the video stream content from the high priority group if it is determined that there is a waste of resources, to filter out the dynamic content streams that truly need high bitrate, and generate a high priority list. If it is determined that there is no waste of resources, the current bitrate allocation scheme is maintained. The configuration equalization module is used to reallocate bandwidth resources according to the high priority list to form a transmission quality configuration; The scheduling and sorting module is used to calculate the dynamic weight of importance for all video streams according to the transmission quality configuration, and arrange the transmission requests in descending order of the dynamic weight to determine the parallel transmission order.
[0063] It should be noted that the video stream parallel processing system for multiple users provided in this embodiment of the invention is used to execute all the process steps of the video stream parallel processing method for multiple users described in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.
[0064] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0065] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for parallel processing of video streams for multiple users, characterized in that, include: Acquire multiple video streams, extract quantitative features of content changes and motion intensity from each video stream, and generate complexity vectors for each video stream. Based on the complexity vector, K-means clustering analysis is performed on all video streams to group video streams with similar complexity into the same group, thus obtaining a set of grouping categories for the video streams; For each group in the grouping category set, obtain the average complexity and total system bandwidth, and analyze the importance weight of each group based on the average complexity and total system bandwidth; Identify high-priority groups whose importance weights exceed a preset threshold, and adjust the corresponding transmission bitrate according to the real-time network status and coding complexity of the high-priority groups to obtain a bitrate allocation scheme. Monitor the transmission delay and stuttering frequency of key video streams under the bitrate allocation scheme to determine whether there is any waste of resources; If it is determined that there is a waste of resources, then a secondary classification is performed on the high priority group based on the dynamic characteristics of the video stream content to filter out the dynamic content streams that truly require high bitrate and generate a high priority list. If it is determined that there is no waste of resources, then the current bitrate allocation scheme is maintained. Based on the high priority list, bandwidth resources are reallocated to form a transmission quality configuration. Based on the transmission quality configuration, dynamic weights of importance are calculated for all video streams, and transmission requests are arranged in descending order of the dynamic weights to determine the parallel transmission order.
2. The method for parallel processing of video streams for multiple users according to claim 1, characterized in that, The process involves acquiring multiple video streams, extracting quantitative features reflecting content changes and motion intensity from each stream, and generating a complexity vector for each video stream, including: Acquire multiple video streams and calculate the grayscale difference between adjacent sampled frames, and generate inter-frame difference data based on the grayscale difference; Macroblock matching is performed on the inter-frame difference data to extract displacement vectors, and the sum of the magnitudes of the displacement vectors is calculated to generate motion intensity data. Video frames are processed using edge detection algorithms to extract image edge gradient information; The motion intensity data and the image edge gradient information are input together into a preset feature mapping model to extract quantitative feature values that reflect spatiotemporal complexity. The deviation is calculated based on the statistical distribution characteristics of the quantized feature values, and the deviation is normalized to obtain the complexity vector of each video stream.
3. The method for parallel processing of video streams for multiple users according to claim 1, characterized in that, Based on the complexity vector, K-means clustering analysis is performed on all video streams to group video streams with similar complexity into the same group, resulting in a set of video stream grouping categories, including: Select initial cluster centroids from the complexity vector; Calculate the Euclidean distance between each complexity vector and the initial cluster centroid, and assign each complexity vector to a temporary group with the minimum Euclidean distance; Update the cluster centroids based on the mean of the complexity vectors within each temporary group, and iterate the assignment and update operations until the centroid changes converge. A group index table is established based on the converged clustering relationships to obtain the set of group categories for the video stream.
4. The method for parallel processing of video streams for multiple users according to claim 1, characterized in that, For each group in the grouping category set, the average complexity and total system bandwidth are obtained. Based on the average complexity and total system bandwidth, the importance weight of each group is analyzed, including: The average motion vector magnitude and average edge pixel density of each video stream group are obtained as the average complexity. The average motion vector magnitude and the average edge pixel density are weighted and fused to obtain the overall complexity. Calculate the ratio of the overall complexity to the current total system bandwidth capacity to obtain the relative bandwidth requirement percentage; If the relative bandwidth demand ratio exceeds the preset congestion warning threshold, the relative bandwidth demand ratio is non-linearly amplified to determine the initial importance weight value. Based on the preliminary importance weight values, the video streams are sorted from high to low to determine their priority order.
5. The method for parallel processing of video streams for multiple users according to claim 1, characterized in that, The process of identifying high-priority groups whose importance weights exceed a preset threshold, and adjusting the corresponding transmission bitrate based on the real-time network status and coding complexity of the high-priority groups to obtain a bitrate allocation scheme includes: If the importance weight of a certain video stream group exceeds a preset importance threshold, it is determined to be a high-priority group; Collect real-time buffer occupancy and packet loss rate data of the high-priority group and calculate the network congestion status value; Obtain the average motion intensity and texture complexity of the high-priority group, and calculate the encoding complexity based on the average motion intensity and texture complexity; The network congestion status value and the coding complexity of the high priority group are input into a preset nonlinear mapping function to obtain the incremental allocation coefficient; The incremental allocation coefficient is superimposed on the preset basic allocation ratio to generate the corrected allocation ratio; The target transmission bitrate is calculated based on the corrected allocation ratio, and the quantization parameters are derived based on the target transmission bitrate to obtain the bitrate allocation scheme.
6. The method for parallel processing of video streams for multiple users according to claim 1, characterized in that, The monitoring of transmission latency and stuttering frequency of key video streams under the bitrate allocation scheme to determine whether there is resource waste includes: Obtain actual transmission delay data and stuttering frequency data for key video streams; The actual transmission delay data and the stuttering frequency data are mapped to a preset quality saturation baseline to identify oversaturated time segments where the transmission quality is higher than the baseline. Calculate the resource overflow magnitude within the oversaturation time segment. If the resource overflow magnitude is greater than a preset redundancy tolerance value, it is determined that there is resource waste; otherwise, it is determined that there is no resource waste.
7. The method for parallel processing of video streams for multiple users according to claim 1, characterized in that, If resource waste is determined to exist, a secondary classification is performed from the high-priority group based on the dynamic characteristics of the video stream content to filter out the dynamic content streams that truly require high bitrates, generating a high-priority list. If no resource waste is determined to exist, the current bitrate allocation scheme is maintained, including: If it is determined that there is a waste of resources, then multi-dimensional spatiotemporal feature data of each video stream is extracted from the high-priority group; Based on the multidimensional spatiotemporal feature data, a feature matrix to be classified is constructed. The feature matrix to be classified is input into a preset support vector machine model to calculate the geometric interval value. Based on the geometric interval value, vectors located on the positive side of the classification hyperplane boundary are selected, and the video streams corresponding to the selected vectors are marked as dynamic content streams that truly require high bitrate. Based on all the marked dynamic content streams, the high priority list is generated. If it is determined that there is no waste of resources, the bandwidth allocation ratio and transmission bitrate of each video stream group will remain unchanged, and the current bitrate allocation scheme will be maintained.
8. The method according to claim 1, characterized in that, The step of reallocating bandwidth resources according to the high-priority list to form a transmission quality configuration includes: The remaining video stream objects are filtered according to the high-priority list; Obtain the motion intensity and texture complexity of the remaining video stream object, calculate the encoding complexity based on the motion intensity and texture complexity, and calculate the lower limit of the bitrate based on the encoding complexity; Calculate the difference between the real-time transmission rate of each remaining video stream object and the corresponding lower limit of the bitrate, and sum them up to obtain the total amount of redundant bandwidth resources; The total amount of redundant bandwidth resources is allocated to the video streams in the high-priority list to obtain the enhanced bitrate value of the key streams; The lower bitrate value and the enhanced bitrate value are combined to form a transmission quality configuration.
9. The method according to claim 1, characterized in that, The step of calculating dynamic weights of importance for all video streams based on the transmission quality configuration, and arranging transmission requests in descending order of the dynamic weights to determine the parallel transmission order includes: Extract the identifiers of each video stream from the transmission quality configuration, and calculate the dynamic importance weights based on the transmission deadline constraints and buffer occupancy rates corresponding to each identifier. The dynamic importance weights of each video stream are mapped to a preset dependency graph, and the logical dependency paths between video streams are determined through topology analysis. The scheduling difficulty coefficient is calculated based on the logical dependency path. If the scheduling difficulty coefficient is higher than the preset difficulty threshold, the logical dependency path is decoupled and a sorted index table is generated. If the scheduling difficulty coefficient is not higher than the preset difficulty threshold, a sorted index table is generated based on the logical dependency path and the dynamic importance weight. The transmission request queues of all video streams are rearranged according to the sorting index table to obtain an ordered queue. The ordered queue is then traversed to determine the final parallel transmission order.
10. A system for parallel processing of video streams for multiple users, characterized in that: include: The feature extraction module is used to acquire multiple video streams, extract quantitative features of each video stream reflecting content changes and motion intensity, and generate complexity vectors for each video stream. The clustering and grouping module is used to perform K-means clustering analysis on all video streams based on the complexity vector, grouping video streams with similar complexity into the same group to obtain a set of grouping categories for the video streams; The weight calculation module is used to obtain the average complexity and total system bandwidth for each group in the grouping category set, and to analyze the importance weight of each group based on the average complexity and total system bandwidth. The bitrate optimization module is used to identify high-priority groups whose importance weight exceeds a preset threshold, and adjust the corresponding transmission bitrate according to the real-time network status and encoding complexity of the high-priority groups to obtain a bitrate allocation scheme. The resource monitoring module is used to monitor the transmission delay and stuttering frequency of key video streams under the bitrate allocation scheme, and to determine whether there is any waste of resources. The list refinement module is used to perform secondary classification based on the dynamic characteristics of the video stream content from the high priority group if it is determined that there is a waste of resources, to filter out the dynamic content streams that truly need high bitrate, and generate a high priority list. If it is determined that there is no waste of resources, the current bitrate allocation scheme is maintained. The configuration equalization module is used to reallocate bandwidth resources according to the high priority list to form a transmission quality configuration; The scheduling and sorting module is used to calculate the dynamic weight of importance for all video streams according to the transmission quality configuration, and arrange the transmission requests in descending order of the dynamic weight to determine the parallel transmission order.