Network-adaptive video conferencing transmission optimization method and system

By analyzing the network status historical data and the current video frame type, and using a fusion model to dynamically adjust the video transmission parameters, the poor adaptability and unreasonable resource allocation of network adaptive video conferencing transmission optimization scheme in the existing technology are solved, and the forward-looking guarantee of the smoothness and clarity of video conferencing is achieved, and the user experience is improved.

CN120358347BActive Publication Date: 2025-08-29SHENZHEN MINRRAY IND CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510830851.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-08-29
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The existing network adaptive video conferencing transmission optimization solution has insufficient accuracy in network status prediction and fails to fully consider the characteristics of video content, resulting in poor adaptability and unreasonable resource allocation, and the inability to ensure the smoothness and clarity of video conferencing in a complex and changeable network environment.

Method used

By systematically analyzing the historical data sequence of network state, deep trends and patterns of network behavior are extracted, and combined with the type of current conference video frames, a fusion model is used for interactive perception and collaborative processing, and video transmission parameters are dynamically adjusted, including code rate, resolution and frame rate, to generate a target code rate that adapts to future network fluctuations and meets current picture quality requirements.

Benefits of technology

In a complex and changeable network environment, the forward-looking guarantee of the fluency and clarity of video conferences is achieved, which improves the user experience, and avoids the poor adaptability and unreasonable resource allocation caused by simply relying on instantaneous network status or ignoring the differences in video content in traditional solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358347B_ABST
    Figure CN120358347B_ABST
Patent Text Reader

Abstract

The present application relates to the field of network communications, and specifically discloses a network-adaptive video conferencing transmission optimization method and system, which systematically analyzes the historical data sequence of the network status to extract the deep trends and patterns of network behavior. At the same time, the different requirements of the current conference video frame type for transmission parameters (especially bit rate) are fully considered, and this key information is integrated into the decision-making process. Through a specially designed fusion model, the historical evolution trend of the network status and the specific requirements of the current video content can be intelligently interactively perceived and collaboratively processed, thereby generating a video target bit rate that can adapt to future network fluctuations and meet current picture quality requirements. Based on this dynamically optimized target bit rate, the system further determines the appropriate target resolution and target frame rate to encode and transmit the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network communications, and more specifically, to a network-adaptive video conferencing transmission optimization method and system. Background Art

[0002] With the rapid development of information technology and the increasing prevalence of global collaboration, video conferencing systems have become an indispensable tool for a variety of scenarios, including corporate communication, distance education, and online healthcare. A high-quality, smooth video conferencing experience is crucial for ensuring efficient communication and user satisfaction. However, the complexity and instability of internet connections, such as bandwidth fluctuations, network congestion, packet loss, and transmission delays, often pose severe challenges to video conferencing transmission quality. This can lead to issues such as screen freezes, pixelation, and audio / video asynchrony, severely impacting the user experience.

[0003] Currently, several network-adaptive video conferencing transmission optimization solutions exist. These typically adjust the video encoding bitrate, resolution, or frame rate by monitoring network parameters (such as bandwidth, latency, and packet loss rate). For example, some solutions may employ rule-based control logic, reducing the bitrate when network deterioration is detected and then attempting to increase it again when the network improves. Other solutions utilize traditional control theory models or simple statistical predictions to guide parameter adjustments. However, these existing solutions often suffer from several drawbacks: First, they primarily focus on instantaneous or short-term changes in network conditions, insufficiently exploring historical trends and complex dynamic characteristics. This results in inaccurate predictions and potentially delayed or volatile adaptive adjustments. Second, their decision-making process fails to fully consider the inherent characteristics of video content. For example, different video frame types (such as I-frames, P-frames, and B-frames, or more macroscopic content complexity, such as static presentations versus dynamic scenes) have varying requirements for bitrate, resolution, and frame rate. Simply applying a one-size-fits-all approach based on network conditions fails to maximize visual quality for specific content while ensuring smoothness, or it wastes unnecessary bandwidth for simple content.

[0004] Therefore, a network adaptive video conferencing transmission optimization solution is desired that can dynamically adjust transmission parameters according to real-time network conditions to optimize the video conferencing transmission effect. Summary of the Invention

[0005] To address the above-mentioned technical problems, the present application is proposed. Embodiments of this application provide a network-adaptive video conferencing transmission optimization method and system. This method systematically analyzes historical data sequences of network status to extract underlying trends and patterns in network behavior. It also fully considers the varying transmission parameter requirements (particularly bitrate) for each type of video frame in the current conference, incorporating this critical information into the decision-making process. Through a specially designed fusion model, the system intelligently interacts and coordinates the historical evolution of network status with the specific requirements of current video content, generating a target video bitrate that can adapt to future network fluctuations while meeting current image quality requirements. Based on this dynamically optimized target bitrate, the system further determines the appropriate target resolution and frame rate for video encoding and transmission. This method effectively overcomes the poor adaptability and irrational resource allocation inherent in traditional solutions that rely solely on instantaneous network status or ignore differences in video content. It aims to proactively ensure the smoothness and clarity of video conferencing in complex and changing network environments, enhancing the user experience.

[0006] According to one aspect of the present application, a network-adaptive video conferencing transmission optimization method is provided, which includes:

[0007] Obtain the time collection of historical network status data from the sender;

[0008] Get the type of the current conference video frame;

[0009] Preprocessing the time set of the network status historical data to obtain a sequence of network status historical data;

[0010] Determining a target video bit rate based on the type of the current conference video frame and the sequence of the network status historical data;

[0011] Determining a target resolution and a target frame rate based on the target video bit rate and the type of the current conference video frame;

[0012] Encoding the current conference video frame based on the target resolution and the target frame rate to obtain an adaptively optimized current conference video frame;

[0013] The adaptively optimized current conference video frame is transmitted for data transmission.

[0014] According to another aspect of the present application, a network-adaptive video conferencing transmission optimization system is provided, comprising:

[0015] The network status data acquisition module is used to obtain the time collection of network status historical data from the sending end;

[0016] The current conference video frame type acquisition module is used to obtain the type of the current conference video frame;

[0017] A network status data preprocessing module, configured to preprocess the time set of the network status historical data to obtain a sequence of network status historical data;

[0018] A video target bit rate generation module is configured to always determine a video target bit rate based on the type of the current conference video frame and the sequence of the network status historical data;

[0019] A target resolution and frame rate determination module, configured to determine a target resolution and a target frame rate based on the target video bit rate and the type of the current conference video frame;

[0020] A conference video frame encoding optimization module for encoding the current conference video frame based on the target resolution and the target frame rate to obtain an adaptively optimized current conference video frame;

[0021] The conference video frame transmission module is used to transmit the current conference video frame after the adaptive optimization.

[0022] Compared to existing technologies, this application provides a network-adaptive video conferencing transmission optimization method and system. This method systematically analyzes historical data sequences of network status to extract underlying trends and patterns in network behavior. It also fully considers the varying transmission parameter requirements (particularly bitrate) for the current conference video frame type, incorporating this critical information into the decision-making process. Through a specially designed fusion model, it intelligently interactively perceives and coordinates the historical evolution of network status with the specific requirements of current video content, generating a target video bitrate that adapts to future network fluctuations while meeting current image quality requirements. Based on this dynamically optimized target bitrate, the system further determines the appropriate target resolution and frame rate for encoding and transmission. This method effectively overcomes the poor adaptability and irrational resource allocation inherent in traditional solutions that rely solely on instantaneous network status or ignore differences in video content. It aims to proactively ensure the smoothness and clarity of video conferencing in complex and changing network environments, enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0024] Figure 1Flowchart of a network-adaptive video conferencing transmission optimization method according to an embodiment of the present application;

[0025] Figure 2 A data flow diagram of a network-adaptive video conferencing transmission optimization method according to an embodiment of the present application;

[0026] Figure 3 A flowchart of determining a target video bit rate based on the type of the current conference video frame and the sequence of the network status historical data according to the network adaptive video conference transmission optimization method of an embodiment of the present application;

[0027] Figure 4 A flowchart of a method for optimizing network-adaptive video conferencing transmission according to an embodiment of the present application for obtaining a network state-video conference type progressive interaction feature vector by passing the network state historical data time series correlation feature vector and the current conference video frame type one-hot encoding vector through a network state-video type progressive complementary interaction perception network;

[0028] Figure 5 This is a block diagram of a network-adaptive video conferencing transmission optimization system according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0030] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0031] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.

[0032] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0033] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0034] To address the challenges of existing solutions, which suffer from inaccurate network status predictions and insufficient awareness of video content characteristics, this solution goes beyond simply responding to network status immediately. Instead, it systematically analyzes historical network status data to extract underlying trends and patterns in network behavior. Furthermore, it fully considers the varying transmission parameter requirements (particularly bitrate) for the current conference video frame type, incorporating this critical information into the decision-making process. A specially designed fusion model intelligently integrates historical network status trends with the specific requirements of current video content, generating a target bitrate that adapts to future network fluctuations while meeting current image quality requirements. Based on this dynamically optimized target bitrate, the system further determines the appropriate target resolution and frame rate for video encoding and transmission. This approach effectively overcomes the poor adaptability and irrational resource allocation inherent in traditional solutions, which rely solely on instantaneous network status or ignore video content variations. It aims to proactively ensure smooth and clear video conferencing in complex and volatile network environments, significantly improving the user experience.

[0035] In the technical solution of the present application, a network-adaptive video conferencing transmission optimization method is proposed. Figure 1 The present invention is a flowchart of a network-adaptive video conferencing transmission optimization method according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of the network adaptive video conferencing transmission optimization method according to the embodiment of the present application. Figure 1 and Figure 2 As shown, the network adaptive video conference transmission optimization method according to the embodiment of the present application includes the steps of: S100, obtaining a time set of network status historical data from a sending end; S200, obtaining the type of a current conference video frame; S300, preprocessing the time set of the network status historical data to obtain a sequence of network status historical data; S400, determining a video target bit rate based on the type of the current conference video frame and the sequence of the network status historical data; S500, determining a target resolution and a target frame rate based on the video target bit rate and the type of the current conference video frame; S600, encoding the current conference video frame based on the target resolution and the target frame rate to obtain an adaptively optimized current conference video frame; S700, transmitting the adaptively optimized current conference video frame for data transmission.

[0036] Specifically, in steps S100 and S200, a time series of historical network status data is obtained from the transmitter, along with the type of the current conference video frame. This historical network status data includes the mean RTT, packet loss rate, receiving bitrate, and sending bitrate. It should be understood that obtaining the time series of historical network status data from the transmitter and the type of the current conference video frame are fundamental and prerequisite for achieving precise dynamic adjustments in network-adaptive video conferencing transmission optimization. These two pieces of information are necessary because historical network status data can reveal the dynamic characteristics and changing trends of network links, such as long-term trends and cyclical fluctuations in network behavior. This provides a more reliable basis for predicting future network conditions, thereby making bitrate decisions more forward-looking. The type of the current conference video frame directly relates to the importance of the frame data and the characteristics of the compression encoding. It reflects the amount of information and compression requirements for the current image. For example, key frames (I frames) typically require a higher bitrate to ensure image quality, while predicted frames (P / B frames) require a relatively lower bitrate. Judging solely by the instantaneous network status is susceptible to interference from short-term factors such as network jitter, leading to delayed or overreacting adaptive adjustments. Ignoring the type of video frame makes it impossible to specifically optimize the transmission quality of key frames or effectively save bandwidth when content complexity is low. Therefore, combining historical network trends with current content characteristics can achieve the unity of network adaptability and content awareness. While ensuring network transmission stability, bitrate resources are dynamically allocated based on content characteristics, thus continuously providing a better video conferencing experience in complex and changing network environments.

[0037] More specifically, in one specific example of this application, first, the transmitter acquires a temporal collection of historical network status data. This typically relies on a continuous network parameter monitoring and recording mechanism. The transmitter periodically collects or obtains a series of key network metrics through feedback from the receiver (e.g., via the RTCP protocol), such as the average round-trip time (RTT), packet loss rate, the actual bitrate received by the receiver, and the transmitter's own bitrate. These data points, including timestamps, are recorded and stored in a buffer or database with a specific time window, forming a chronologically ordered data set. For example, the system can set a sliding time window, such as the past few seconds or tens of seconds, and continuously update the sequence of network status parameters within this window. When a decision is needed, relevant data is extracted from this historical data set. Second, the type of the current conference video frame is typically determined during the video encoding process or in the pre-encoding processing stage. When processing each frame of video, the video encoder classifies it into different types based on its encoding strategy (such as GOP structure settings and scene change detection results). Common types include independently encoded key frames (I frames), forward-predicted P frames, or bidirectionally predicted B frames. When a video frame is about to be processed for adaptive optimization decisions, its type information can be obtained directly from the video encoder or video processing module. This information is crucial for subsequently determining the target bitrate for the frame, as different types of frames have very different requirements for image quality assurance and compression rate. In this way, the sender can simultaneously understand the historical performance of the network link and the basic properties of the current video content to be transmitted, laying a solid data foundation for subsequent intelligent determination of the target bitrate, resolution, and frame rate based on this information.

[0038] Specifically, in step S300, the time collection of historical network status data is preprocessed to obtain a sequence of historical network status data. It should be understood that the originally collected historical network status data, such as mean RTT, packet loss rate, received bit rate, and sent bit rate, often have different physical units, dimensions, and numerical ranges. For example, RTT may be measured in milliseconds, with values ​​fluctuating between tens and hundreds, while packet loss rate is a percentage, with values ​​between 0 and 100 (or 0 to 1). If this heterogeneous data is directly input into subsequent complex models without processing, features with a larger numerical range may disproportionately dominate the model's learning process, obscuring the contributions of other, smaller but equally important features. This can lead to difficulties in model training, slow convergence, and even the inability to accurately capture the true relationships between features. Furthermore, these data points occur over time, and their inherent temporal dependencies are crucial for predicting future network trends. Therefore, the original "time collection of historical network status data" must be converted into a structured "sequence" that clearly reflects these temporal relationships before it can be effectively utilized by the time series model.

[0039] Specifically, in an embodiment of the present application, the time collection of network status historical data is preprocessed to obtain a sequence of network status historical data, including normalizing and serializing the time collection of network status historical data to obtain the sequence of network status historical data. Specifically, first, normalization allows network status data of different dimensions (such as RTT and packet loss rate) to be uniformly mapped to similar numerical intervals, such as the commonly used intervals [0, 1] or [-1, 1]. This eliminates the dimensionality effects and value range differences between different features, allowing the model to treat all input features fairly and avoiding learning bias caused by numerical differences. This accelerates model training convergence and improves model stability and generalization. Second, serialization organizes what might originally be a collection of data points with timestamps into a strictly chronologically ordered time series. This provides structured input for the subsequent extraction of temporal correlation features of the network status, ensuring that the model can effectively learn and capture key information such as dynamic patterns, periodic patterns, and sudden changes in the network status over time.

[0040] Specifically, in step S400, a target video bitrate is determined based on the type of the current conference video frame and the sequence of historical network status data. It should be understood that traditional rate control strategies often focus on responding to the current instantaneous network state or fail to fully distinguish the characteristics of different video content. This leads to problems such as insufficient adaptability, inaccurate predictions, and suboptimal resource allocation. Furthermore, the sequence of historical network status data contains in-depth information about network behavior, such as trends, periodicity, and bursts. It can reveal the true carrying capacity of a network link and potential future fluctuations far more than a single instantaneous data point. Furthermore, the bitrate requirements for the current conference video frame type, such as key frames (I-frames) with high information content and high quality requirements, and predicted frames (P-frames or B-frames) with relatively low information content and high compressibility, vary significantly. Therefore, effectively combining these two dimensions—the long-term dynamic characteristics of the network and the real-time encoding requirements of the video content—can overcome the limitations of traditional solutions. Specifically, through in-depth analysis of historical network conditions and keen awareness of the characteristics of current video content, the system dynamically calculates an "optimal" target video bitrate that adapts to current and foreseeable future network conditions while meeting specific video frame quality requirements. This "optimal" goal isn't simply about pursuing the highest bitrate possible, but rather about allocating appropriate bitrate resources to video frames of varying importance and compression characteristics while ensuring smooth transmission (i.e., network capacity). This approach aims to achieve the optimal balance between video clarity and smoothness in a volatile network environment. This precisely determined target bitrate serves as a key basis for subsequent adjustments to video resolution and frame rate, ultimately guiding the video encoding process.

[0041] Figure 3 The present invention is a flowchart of determining the target resolution and target frame rate based on the target video bit rate and the type of the current conference video frame according to the network adaptive video conference transmission optimization method of the embodiment of the present application. Figure 3 As shown, according to the network adaptive video conferencing transmission optimization method of the embodiment of the present application, step S400 includes: S410, arranging the sequence of the network status historical data into a network status historical data time series matrix according to the network status sample dimension and the time dimension; S420, extracting the network status time series association features from the network status historical data time series matrix to obtain the network status historical data time series association feature vector; S430, performing one-hot encoding on the type of the current conference video frame to obtain the current conference video frame type one-hot encoding vector; S440, passing the network status historical data time series association feature vector and the current conference video frame type one-hot encoding vector through a network status-video type progressive complementary interaction perception network to obtain a network status-video conference type progressive interaction feature vector; S450, feature decoding the network status-video conference type progressive interaction feature vector to obtain the video target bit rate.

[0042] Specifically, in step S410, the sequence of network state historical data is organized into a network state historical data time series matrix according to the network state sample dimension and the time dimension. It should be understood that while the "sequence of network state historical data" obtained through the previous preprocessing has already solved the fundamental problems of data heterogeneity and time series, for many advanced analytical models, especially deep learning models designed to capture complex time dependencies and multivariate interactions, directly using a one-dimensional sequence or simple list structure may not be efficient or fully utilize the inherent structure of the data. Network state itself is multidimensional (the "network state sample dimension," for example, multiple metrics such as RTT, packet loss rate, transmission bit rate, and reception bit rate), and these multidimensional metrics evolve over time (the "time dimension"). Therefore, in the technical solution of the present application, the sequence of network state historical data is further organized into a network state historical data time series matrix according to the network state sample dimension and the time dimension. It is worth noting that organizing these multidimensional network state historical time series data into a matrix, with one axis representing time steps and the other representing different network state parameters, can present the historical dynamics of the network in a highly structured manner. This time-series matrix of historical network state data clearly displays the specific values ​​of multiple network parameters at consecutive time points and their interrelationships. This matrix format not only facilitates subsequent mathematical operations and feature transformations, but also fully preserves and explicitly expresses the distribution of network state information on the two-dimensional "feature-time" plane. This provides an ideal data foundation for further exploration of potential, nonlinear, time-series correlations in network state, such as periodic fluctuations, trend changes, and sudden congestion patterns.

[0043] Specifically, in step S420, network state temporal correlation features are extracted from the network state historical data time series matrix to obtain a network state historical data time series correlation feature vector. It should be understood that although the original network state historical data time series matrix structuredly presents the changes in multi-dimensional network parameters over time, it is still a relatively primitive data representation. Directly using the entire high-dimensional time series matrix for subsequent decision fusion is not only computationally expensive but may also contain redundant information or noise. More importantly, it fails to explicitly reveal the deep, nonlinear correlation patterns and long-term dependencies between different time points and different network parameters, which are the key to accurately judging network trends and predicting future states. Traditional statistical methods or simple time series analysis are difficult to fully explore these complex internal connections. Therefore, in the technical solution of the present application, network state temporal correlation features are extracted from the network state historical data time series matrix to obtain a network state historical data time series correlation feature vector. The high-dimensional network state historical data time series matrix is ​​mapped and condensed into a low-dimensional but more informative network state historical data time series correlation feature vector. This feature vector is designed to encapsulate and characterize the most critical and representative temporal dynamic characteristics in historical network data, such as the evolution pattern of network congestion, the periodicity of bandwidth fluctuations, and the potential causal or accompanying relationships between different parameters (such as RTT and packet loss rate). Because the dilated convolutional neural network can expand the receptive field by adjusting the dilation rate, thereby effectively capturing long-term dependencies in time series without increasing excessive computational complexity and parameters, it is very suitable for learning these complex correlation features from network time series data. Therefore, in an embodiment of the present application, network state time series correlation features are extracted from the network state historical data time series matrix to obtain a network state historical data time series correlation feature vector, including: passing the network state historical data time series matrix through a network state time series correlation feature extractor based on a dilated convolutional neural network model to obtain the network state historical data time series correlation feature vector. The goal is to generate a compact feature representation that can accurately reflect the overall historical status of the network and highlight key change trends, providing a high-quality, refined network state "portrait" for subsequent intelligent fusion with video frame type information. This time-series correlation feature vector of historical network state data captures the essence of network dynamics more effectively than raw time series data or simple statistical features. This feature vector is no longer simply a list of data; it abstracts and summarizes network behavior patterns, eliminating noise and redundancy and enhancing key information. This enables subsequent fusion with video frame type features during the progressive, complementary interactive perception of network state and video type, enabling a more precise understanding of historical network trends, leading to a more reasonable and forward-looking target video bitrate.

[0044] Specifically, in step S430, the type of the current conference video frame is one-hot encoded to obtain a one-hot encoded vector of the current conference video frame type. It should be understood that subsequent intelligent decision-making models usually require numerical input rather than original category labels (such as "I frame", "P frame", "B frame"). If these category labels are simply replaced with arbitrary numbers (for example, I=1, P=2, B=3), an ordinal relationship or size comparison that does not exist will be inadvertently introduced, misleading the model into believing that P frames are "greater than" I frames in some sense. This is inconsistent with the actual characteristics of the video frame type and will interfere with the model's learning process and decision accuracy. One-hot encoding can convert these discrete, unordered category features into a numerical, high-dimensional sparse vector form that is easy for the model to understand and process, avoiding this artificially introduced bias. Therefore, in the technical solution of the present application, the type of the current conference video frame is further one-hot encoded to obtain a one-hot encoded vector of the current conference video frame type. In this way, the important content feature of the type of the current conference video frame can be represented in a way that is friendly and unbiased to the machine learning model. Through one-hot encoding, each video frame type (such as I-frames, P-frames, and B-frames) is mapped to a new vector space with dimensions equal to the total number of frame types, where the dimension corresponding to the current frame type is 1 and the remaining dimensions are 0. For example, if there are three frame types, I-frames might be encoded as [1, 0, 0], P-frames as [0, 1, 0], and B-frames as [0, 0, 1]. This ensures that different frame types are input to the model as independent and equal category identifiers, without any numerical size or order implications, accurately reflecting their essential differences as distinct content units. This one-hot encoded vector is then used as input along with feature vectors extracted from historical network state data for subsequent fusion analysis by the interactive perception network.

[0045] Specifically, in step S440, the network state historical data temporal correlation feature vector and the current conference video frame type one-hot encoding vector are applied to a network state-video type progressive complementary interaction perception network to obtain a network state-video conference type progressive interaction feature vector. It should be understood that simply concatenating or linearly combining the extracted network state historical temporal correlation features with the current conference video frame type features is far from sufficient to capture the profound and nonlinear dependency between the two, which is crucial for optimizing video conferencing transmission. Dynamic changes in the network environment (represented by the network state historical data temporal correlation feature vector) have varying impacts on the transmission quality of different video frame types (represented by the current conference video frame type one-hot encoding vector). Conversely, different video frames also have different requirements for network resources. For example, a stable, high-bandwidth network environment is extremely beneficial for critical I-frames, allowing for a high bitrate to be allocated to ensure clarity. However, a highly volatile network, even if the current instantaneous state is acceptable, may require a more conservative bitrate strategy for upcoming P-frames or B-frames. This complex, multi-layered interactive judgment cannot be effectively modeled using a simple model. Therefore, in the technical solution of this application, the network status historical data temporal correlation feature vector and the current conference video frame type one-hot encoding vector are further processed through a network status-video type progressive complementary interaction perception network to obtain a network status-video conference type progressive interaction feature vector. Through the processing of the network status-video type progressive complementary interaction perception network, the network status and video frame type are not only treated as independent factors, but their synergistic effects and complementary information at different levels of abstraction are deeply explored. The network uses its multi-level, staged processing paradigm—first, independently extracting multi-level implicit features from the network status historical data temporal correlation feature vector and the current conference video frame type one-hot encoding vector (generating their respective low-level, mid-level, and deep-level implicit features), then performing feature fusion at the corresponding abstraction levels (the low level focuses on detail correspondence, the mid-level focuses on structural association, and the deep level focuses on semantic synergy). Finally, the interaction patterns captured at these different levels are integrated through a progressive complementary perception fusion mechanism. This final "network state-video conference type progressive interaction feature vector" aims to generate a more comprehensive and robust joint representation. It fully reflects the deep coupling between current network trends and current video content requirements, providing high-quality, high-information decision-making for downstream target video bitrate determination. Through this progressive and complementary interactive perception, the system transcends simple response to a single factor and achieves a true integration of "content awareness" and "network prediction."For example, the network can understand that even if the deep features of the network's historical state indicate stability (the result of deep fusion), if the current frame is a P frame (the result of interaction between low- and mid-level features) and has recently experienced slight jitter (the result of interaction between mid-level features), the final interactive feature vector will guide a compromise bitrate that leverages network stability while taking into account the characteristics of P frames and recent minor fluctuations. This effectively avoids inaccurate adjustments and underutilized resources caused by ignoring historical network trends and the influence of video frame types, ultimately ensuring the highest possible clarity and smoothness of video conferences in complex and changing network environments.

[0046] Figure 4 This is a flowchart of the network adaptive video conferencing transmission optimization method according to an embodiment of the present application, which passes the sequence of the audio mel-spectrogram semantic feature vector and the video stream segment semantic feature vector through a video segment semantic feature search module based on graph learning to obtain a video stream segment-audio semantic search response encoding vector. Figure 4 As shown, according to the network adaptive video conferencing transmission optimization method of the embodiment of the present application, step S440 includes: S441, performing multi-level implicit feature extraction on the network status historical data time series association feature vector and the current conference video frame type one-hot encoding vector to obtain the network status historical data time series middle-level implicit feature encoding vector, the current conference video frame type middle-level implicit feature encoding vector, the network status historical data time series deep-level implicit feature encoding vector and the current conference video frame type deep-level implicit feature encoding vector; S442, based on the network status historical data time series middle-level implicit feature encoding vector, the current conference video frame type middle-level implicit feature encoding vector, the network status historical data time series deep-level implicit feature encoding vector and the current conference video frame type deep-level implicit feature encoding vector, performing multi-level feature joint perception on the network status historical data time series association feature vector and the current conference video frame type one-hot encoding vector to obtain the network status-video conference type progressive interaction feature vector.

[0047] More specifically, step S441 performs multi-level implicit feature extraction on the network status history data time series associated feature vector and the current conference video frame type one-hot encoding vector to obtain the network status history data time series middle-layer implicit feature encoding vector, the current conference video frame type middle-layer implicit feature encoding vector, the network status history data time series deep-layer implicit feature encoding vector, and the current conference video frame type deep-layer implicit feature encoding vector, which can be expressed as follows:

[0048]

[0049]

[0050]

[0051]

[0052] in, is the time series correlation feature vector of the network status historical data, and are the network status historical data time series trainable middle-layer weight matrix and the network status historical data time series trainable middle-layer bias vector, is the implicit feature encoding vector of the network status historical data time series, is the one-hot encoding vector of the current conference video frame type, and They are respectively the trainable middle-layer weight matrix and the trainable middle-layer bias vector of the current conference video frame type, is the mid-level implicit feature encoding vector of the current conference video frame type, and They are respectively the network status historical data time series trainable deep weight matrix and the network status historical data time series trainable deep bias vector, is the deep implicit feature encoding vector of the network status historical data time series, and They are respectively the trainable deep weight matrix and the trainable deep bias vector of the current conference video frame type, The deep implicit feature encoding vector of the current conference video frame type.

[0053] It should be understood that directly using the raw temporal correlation feature vectors of historical network state data and the one-hot encoding vectors of the current conference video frame type for a single-level interaction cannot fully capture the complex relationships between the two at different levels of abstraction. The interaction between the temporal characteristics of network state and the content attributes of video frames is not a single-dimensional one. For example, the relationship between the long-term stability of the network and the global importance of frame types (such as I-frames) and the relationship between short-term burst jitter and the local coding requirements of frame types (such as P-frames) are interactions at different levels. Therefore, it is necessary to first perform independent deep nonlinear transformations on each of these two input feature vectors, projecting them into different feature subspaces. This reveals their inherent, deeper properties at multiple levels of abstraction, paving the way for subsequent targeted interactions. Through the hierarchical feature learning capabilities of deep neural networks, the temporal correlation feature vectors of historical network state data are transformed to obtain "intermediate-level latent feature encoding vectors of historical network state data" that represent medium-term network trends and structural fluctuations, and "deep-level latent feature encoding vectors of historical network state data" that represent long-term network dynamics and global semantics. Similarly, the one-hot encoding vector of the current conference video frame type is transformed to obtain the "mid-level implicit feature encoding vector of the current conference video frame type," which reflects the typical requirements of this frame type in moderately complex scenarios (such as within a GOP structure), and the "deep-level implicit feature encoding vector of the current conference video frame type," which reflects its contribution to overall video quality and fundamental encoding requirements. These different levels of implicit feature vectors each encapsulate the core information of the original input at a specific level of abstraction, laying a solid foundation for more precise and targeted feature interaction at the corresponding level.

[0054] Correspondingly, according to an embodiment of the present application, step S442 includes: performing low-level feature fusion on the network status historical data time series associated feature vector and the current conference video frame type one-hot encoding vector to obtain a network status-video conference type low-level fusion feature encoding vector; performing high-level feature fusion on the network status historical data time series middle-level implicit feature encoding vector and the current conference video frame type middle-level implicit feature encoding vector to obtain a network status-video conference type middle-level fusion feature encoding vector and a network status-video conference type deep-level fusion feature encoding vector; performing progressive complementary perception fusion on the network status-video conference type low-level fusion feature encoding vector, the network status-video conference type middle-level fusion feature encoding vector and the network status-video conference type deep-level fusion feature encoding vector to obtain the network status-video conference type progressive interaction feature vector.

[0055] More specifically, a low-level feature fusion is performed on the network status historical data temporal correlation feature vector and the current conference video frame type one-hot encoding vector to obtain a network status-video conference type low-level fusion feature encoding vector, which is expressed as follows:

[0056]

[0057] in, For positional addition, is a multi-layer perceptron, Encode the low-level fusion feature vector of network status and video conference type.

[0058] It's understandable that while mid-level and deep-level feature extraction and fusion will be performed later, these deep abstraction processes inevitably smooth out or lose some of the high-frequency details and most direct correspondences in the original input features. While the temporal correlation feature vectors of historical network state data have been refined, they still contain relatively direct details of network fluctuations. The one-hot encoding vector of the current conference video frame type directly and explicitly indicates the fundamental properties of the current frame. To ensure that the final decision fully utilizes this most raw and direct interaction information and prevent the loss of sensitivity to instantaneous changes or precise matches during subsequent abstraction, it is necessary to perform "early fusion" at the lowest level, as close to the input features as possible. In the context of video conferencing, low-level feature fusion directly captures the direct, unabstracted correlation between the most immediate and subtle fluctuations in the network state history (such as instantaneous spikes in RTT or bursts of packet loss) and the specific type of the current frame: I-frame, P-frame, or B-frame. For example, the interaction between a recent indication of network congestion (high-frequency information) and the current I-frame, which is sensitive to packet loss (precise correspondence), is most directly reflected in low-level fusion. This low-level fusion feature encoding vector is designed to encapsulate this most primitive and fundamental correspondence pattern, providing a perspective that preserves the original details for subsequent progressive complementary perceptual fusion.

[0059] Accordingly, according to an embodiment of the present application, high-level feature fusion is performed on the mid-level implicit feature coding vector of the network status historical data time series and the mid-level implicit feature coding vector of the current conference video frame type to obtain the mid-level fusion feature coding vector of the network status-video conference type and the deep-level fusion feature coding vector of the network status-video conference type, including: mid-level feature fusion is performed on the mid-level implicit feature coding vector of the network status historical data time series and the mid-level implicit feature coding vector of the current conference video frame type to obtain the mid-level fusion feature coding vector of the network status-video conference type; deep feature fusion is performed on the deep-level implicit feature coding vector of the network status historical data time series and the deep-level implicit feature coding vector of the current conference video frame type to obtain the deep-level fusion feature coding vector of the network status-video conference type.

[0060] More specifically, mid-level feature fusion is performed on the mid-level implicit feature coding vector of the network status historical data time series and the mid-level implicit feature coding vector of the current conference video frame type to obtain the mid-level fusion feature coding vector of the network status-video conference type, which is expressed as follows:

[0061] ;

[0062]

[0063] in, For attention fusion processing, is the point product by position, for The scale, for function, Encode the feature vector for the mid-level fusion of network status and video conference type.

[0064] Understandably, while low-level fusion preserves original details, it may not be sufficient to capture larger, more structural patterns. Jumping directly to deep semantic fusion may overlook important intermediate-scale interactions between details and the global picture. The interplay between the historical evolution of network state (e.g., trends and periodic fluctuations over a few seconds to tens of seconds) and video frame type (considering mid-level attributes such as its position in the GOP structure and its dependence on previous and subsequent frames) is often most pronounced at a "structural" level, beyond instantaneous correspondence but short of global semantics. Mid-level features have already filtered out some noise and encode more stable patterns. Therefore, fusion at this level can more effectively identify and exploit the interactions between these initially abstracted and refined stable patterns. Specifically, mid-level latent features of network state represent structural information such as the formation and relief of network congestion and the cyclical fluctuations of bandwidth; whereas mid-level latent features of video frame type reflect its role in the coding unit (e.g., GOP), its dependence on reference frames, and its short- to medium-term contribution to overall visual coherence. The goal of mid-level fusion is to model cross-source structural associations or compositional relationships, capturing these more complex dependencies than those at lower levels. For example, it can learn how to adjust the target bitrate for a P-frame following a key I-frame (a mid-level feature of video frames, reflecting its structural dependency) when the network exhibits a slight deterioration trend lasting several seconds (a mid-level network feature). This judgment is more robust than simply responding to instantaneous network spikes (low-level) and more sensitive than simply considering long-term average network quality (deeper levels).

[0065] More specifically, deep feature fusion is performed on the network status historical data time series deep implicit feature coding vector and the current conference video frame type deep implicit feature coding vector to obtain the network status-video conference type deep level fusion feature coding vector, which is expressed as:

[0066]

[0067]

[0068] in, is the trainable deep weight matrix, for activation function, and is the low-rank projection matrix, for activation function, The network status-video conference type deep level fusion implicit feature encoding vector, Deep-level fusion feature encoding vector for network status-video conference type.

[0069] It's understandable that merely preserving low-level details and capturing mid-level structural associations is insufficient for a holistic, strategic understanding of network status and video content requirements. Long-term trends in network status, overall health (e.g., average bandwidth, packet loss rate, and stability assessments over several minutes or longer), and the fundamental importance of video frame types within the entire video stream, or their core contribution to the perceived quality of the final interactive experience (e.g., the absolute criticality of I-frames as a decoding starting point, or the expendability of B-frames in specific scenarios), all represent "deep semantics" that require a more macro and abstract level of understanding. Therefore, it's necessary to fuse the most abstracted and information-compressed deep features of network status with the deep features of video frame types to model their synergy at the highest level. In video conferencing scenarios, understanding requires questions like, "How can we ensure the basic transmission quality of key I-frames (deep features representing their core semantic value) when the overall network environment is chronically poor (deep network features)?" or "How can we maximize the encoding quality of all frame types to enhance the overall visual experience when the network is chronically stable and high-quality?" This deep fusion aims to extract a highly generalized and guiding interactive judgment that focuses more on long-term strategies rather than short-term tactics, setting a macro tone or constraint for the entire bitrate adaptation system.

[0070] More specifically, according to an embodiment of the present application, the network state-video conference type low-level fusion feature coding vector, the network state-video conference type mid-level fusion feature coding vector, and the network state-video conference type deep-level fusion feature coding vector are subjected to progressive complementary perception fusion to obtain the network state-video conference type progressive interaction feature vector, which is expressed as follows:

[0071]

[0072]

[0073] in, for and cascade operation, is the trainable cascade weight matrix, for function, The shallow-middle level fusion gating adjustment parameters for network status-video conference type are: The shallow-middle level fusion feature encoding vector of network status-video conference type, 、 and are the trainable query matrix, the trainable key matrix, and the trainable value matrix, respectively. 、 and They are the shallow-middle level fusion query feature vector of network status-video conference type, the deep-level fusion key feature vector of network status-video conference type, and the middle-level fusion value feature vector of network status-video conference type. for The scale, is the network status-video conference type progressive interaction feature vector.

[0074] It should be understood that the previously generated low-, medium-, and deep-level fusion feature encoding vectors for network status and video conference type capture the interactive information between network status and video frame type at different granularities (temporal details, structural patterns, and global semantics). However, each of these vectors is independent and only interprets this complex relationship from a single perspective. Simply concatenating or averaging these three layers could lead to information redundancy, dilution of important features, or ineffective resolution of potential conflicts between different layers, potentially resulting in performance degradation. Therefore, a progressive complementary perceptual fusion approach is needed to intelligently integrate these multi-layered interactive information, ensuring that they reinforce rather than interfere with each other. This ultimately constructs a unified representation of the interactive relationship between network status and video content, namely the "progressive interaction feature vector for network status and video conference type," providing a comprehensive, profound, and synergistic understanding. The "progressive" approach here implies an orderly and progressive fusion process, with lower-level information providing a foundation for the middle-level layers, which in turn provide context for the higher-level layers. In other words, integration occurs gradually through a dynamic weighting or gating mechanism. "Complementarity" emphasizes the uniqueness and irreplaceability of features at different levels. Each contributes different aspects of the overall understanding, and the goal of fusion is to leverage their strengths and offset their weaknesses, forming a more powerful and comprehensive feature representation than any single level alone. This fusion dynamically integrates interactive information from different levels of abstraction, enabling the final network state-videoconference type progressive interaction feature vector to simultaneously account for instantaneous changes, medium-term trends, and long-term dynamics between network state and videoconference type, while skillfully integrating it with the immediate needs, structural importance, and core value of the video frame. In videoconference scenarios, this means the system can make extremely accurate and robust judgments based on this feature vector. For example, if a low-level layer indicates a brief network spike, but mid- to deep-level layers indicate overall network stability and the current frame is a critical I-frame, progressive complementary fusion can weigh this information and decide to moderately smooth the impact of the spike, prioritizing I-frame quality rather than blindly and drastically downgrading the bitrate. Conversely, if information from multiple levels indicates network degradation, a more conservative bitrate strategy will be decisively adopted. This intelligent integration capability enables the system to demonstrate excellent adaptability, foresight, and stability when faced with complex, changeable, and uncertain real-world network environments, ultimately significantly improving the smoothness and visual quality of video conferencing and ensuring a high-quality user experience under various network conditions.

[0075] Preferably, when the network state-video conference type low-level fusion feature coding vector, the network state-video conference type middle-level fusion feature coding vector and the network state-video conference type deep-level fusion feature coding vector are subjected to step-by-step complementary perception aggregation, the network state-video conference type low-level fusion feature coding vector and the network status-video conference type mid-level fusion feature encoding vector Essentially generates network status-video conference type shallow-middle level fusion gating adjustment parameters , and then adjust the parameters based on the shallow-middle level fusion gate of the network status-video conference type The phase complementation is used for fusion, and the network state-video conference type deep level fusion feature coding vector is further combined with the network state-video conference type deep level fusion feature coding vector Through the weight matrix 、 and The order parameter field is restricted to facilitate the interaction. Here, the interaction is hindered due to the instability of the bound structure solution of the order parameter field, thereby affecting the characterization performance of the network state-video conference type progressive interaction feature vector.

[0076] Therefore, for the pre-trained matrix 、 and , first calculate the decomposition field path integral representation of the order parameter field of each matrix union:

[0077]

[0078]

[0079]

[0080] in, 、 and All are integral operations. 、 and All for 、 and The order parameter field decomposition field path integral representation matrix, the above multiplications are all matrix dot multiplications.

[0081] Then, the field uniform state alignment loss function is defined as :

[0082]

[0083] in represents the nuclear norm of the matrix, that is, the sum of the eigenvalues ​​of the matrix, is the scaling hyperparameter.

[0084] In this way, the path integral of each decomposition field of the order parameter field is used to weaken the non-integrable superposition of the bound connection of the order parameter field to prevent the occurrence of geometric phase superposition that hinders the interaction of local field structures. Then, the field uniformity equivalence of the sum of eigenvalues ​​can be used to avoid the instability of the bound structure solution, thereby realizing the weight matrix 、 and The interaction is maintained, improving the representation performance of the network state-video conference type progressive interaction feature vector.

[0085] Specifically, in step S500, the target resolution and target frame rate are determined based on the target video bit rate and the type of the current conference video frame. It should be understood that the clarity of the video (mainly determined by the resolution) and the smoothness of the video (mainly determined by the frame rate) are mutually constrained and both consume bit rate. When the total bit rate is limited, a trade-off must be made between the two. Different types of video frames (such as I frames, P frames, and B frames) have significantly different characteristics for the final decoding quality and bit rate consumption. For example, I frames, as key frames, have a significant impact on the quality of subsequent frames and usually require a higher bit rate to ensure their clarity, while B frames have higher fault tolerance. Therefore, the resolution and frame rate are dynamically adjusted according to the target bit rate and the current frame type, aiming to most effectively allocate limited bit rate resources to the dimensions most needed by the current frame, in order to maximize the user's visual experience while meeting transmission constraints.

[0086] More specifically, in one specific example of this application, the first step involves preliminary parameter range definition driven by bitrate. Based on the input target video bitrate, the system references a preset bitrate-resolution-frame rate mapping table or empirical model to initially select a set or a roughly suitable resolution and frame rate range. For example, a very low target bitrate (e.g., a few hundred kbps) typically corresponds to a lower resolution (e.g., 360p or even lower) and potentially a lower frame rate (e.g., 15-20 fps), while a higher target bitrate (e.g., several Mbps) allows for higher resolutions (e.g., 720p or 1080p) and standard frame rates (e.g., 25-30 fps). This step primarily ensures that the selected parameters can theoretically be supported by the target bitrate. The second step involves adjusting parameter biases based on the current conference video frame type. At this stage, the system refines its selection. If the current frame is an I-frame, given its criticality as a decoding reference, the system tends to prioritize a higher resolution within the bitrate budget to ensure basic image clarity. If the bitrate is very tight, a slight sacrifice in frame rate may be made (but only to maintain a minimum acceptable level of smoothness). If the current frame is a P-frame, the system seeks a relatively balanced balance between resolution and frame rate, as P-frames carry the majority of motion information and therefore require a certain degree of smoothness. If the current frame is a B-frame, due to its non-reference and discardable nature, when bitrate is tight, the system will significantly reduce its resolution and even frame rate to save bitrate for more important frames or provide a buffer during network fluctuations. This adjustment may also take into account content characteristics. For example, frame rate may be prioritized for highly dynamic scenes, while resolution may be prioritized for conferences with more static content. The third step is the final decision and constraint checking. Based on the analysis in the first two steps, the system selects a specific target resolution and frame rate combination. This decision may be made by consulting a more refined predefined strategy table (for example, presetting optimal resolution / frame rate pairs for different bitrate levels and frame type combinations) or by using a simple utility function to evaluate the expected perceptual quality of different combinations under the current conditions. At the same time, feasibility checks are performed to ensure that the selected parameter combination meets the capabilities of the encoder and certain minimum quality standards (for example, the frame rate is not lower than a certain threshold to avoid image freezes, and the resolution is not lower than a certain threshold to ensure basic legibility).

[0087] Specifically, in step S600, the current conference video frame is encoded based on the target resolution and the target frame rate to obtain an adaptively optimized current conference video frame. It should be understood that resolution determines image clarity, and frame rate determines motion smoothness. These two parameters are key factors in resource consumption during video encoding. Setting these parameters in advance means that the encoder already knows the spatial dimensions of its output image and the number of frames per unit time when it begins processing each frame. This allows for targeted encoding strategies and bit allocation, ensuring that the generated video stream meets both transmission capabilities and the user's core requirements for clarity and smoothness. Therefore, in the technical solution of the present application, the current conference video frame is encoded based on the target resolution and the target frame rate to obtain an adaptively optimized current conference video frame. This effectively converts limited channel resources (represented by the previously determined "video target bit rate") into user-perceived video quality. Specifically, by adjusting the spatial detail and temporal smoothness of the video stream under dynamically changing network conditions, it strives to achieve an optimal balance between encoding and transmission costs and the final visual experience.

[0088] More specifically, in one specific example of this application, the first step is frame capture and parameter adaptation preprocessing. The system obtains raw video frames from a video capture device (such as a camera). These raw frames typically have the device's native resolution and frame rate. Subsequently, these raw frames must be scaled to match the previously determined "target resolution" and "target frame rate." If the raw resolution is higher than the target resolution, image downsampling (e.g., using bilinear interpolation, bicubic interpolation, or more advanced learning-based scaling algorithms) is required to reduce each frame to the target size. Conversely, if the target resolution is higher (less common in rate-constrained adaptive scenarios, unless there is a specific super-resolution requirement), upsampling may be required. Similarly, with respect to frame rate, if the raw frame rate is higher than the target frame rate, temporal frame extraction is required, such as simply discarding excess frames or employing a more intelligent frame selection strategy to ensure motion consistency. If the target frame rate is lower than the raw frame rate, no special processing is required; frames can simply be delivered at the target rate. (In real-time encoding, frame interpolation to increase the frame rate is generally not used due to its computational complexity and potential artifacts.) This step ensures that the data fed into the core encoder conforms to the "adaptive optimization" directives in both spatial and temporal dimensions. The second step is core video coding compression. Pre-processed video frames with the target resolution and frame rate are fed into the video encoder. Depending on the selected coding standard, the encoder performs a series of complex operations to remove redundant information. These operations primarily include: intra-frame prediction using spatial correlation between pixels within a frame, or inter-frame prediction using temporal correlation between consecutive frames (this process uses the current frame type, such as I, P, or B, to guide the prediction mode selection); transforming the prediction residual (e.g., DCT) to concentrate its energy; quantizing the transform coefficients, which is key to lossy compression and is directly constrained by the target bitrate (the quantization parameter (QP) is dynamically adjusted by the rate control module); and finally, entropy coding (e.g., CABAC or Huffman coding) is performed on the quantized coefficients and other auxiliary information (e.g., motion vectors, prediction mode, etc.) to produce the final compressed bitstream. During the entire encoding process, the encoder's internal bitrate control mechanism will refer to the determined "video target bitrate" and continuously adjust parameters such as the quantization level, striving to make the output bitstream size close to the target value, while ensuring that the video quality is maintained as well as possible at the target resolution and frame rate.

[0089] Specifically, in step S700, the adaptively optimized current conference video frame is transmitted for data transmission. It should be understood that after a complex series of network status perception, video content analysis, target bitrate determination, and adaptive adjustment of encoding parameters (resolution, frame rate) based on this, the generated "adaptively optimized current conference video frame" represents the most optimized visual data packet under the current constraints. Therefore, the adaptively optimized current conference video frame is further transmitted for data transmission. Transmission is the bridge connecting the intelligent decision-making of the sending end and the user experience of the receiving end. Its direct purpose is to efficiently and reliably (within the allowed real-time range) deliver carefully encoded video data to one or more remote participants, allowing them to view the video images of the participants in real time, thereby achieving effective remote communication and collaboration.

[0090] In summary, the network-adaptive video conferencing transmission optimization method according to the embodiments of the present application is described. This method systematically analyzes historical data sequences of network status to extract underlying trends and patterns in network behavior. It also fully considers the varying transmission parameter requirements (particularly bitrate) for the current conference video frame type, incorporating this critical information into the decision-making process. Through a specially designed fusion model, the historical evolution of network status and the specific requirements of current video content are intelligently interactively perceived and collaboratively processed, generating a target video bitrate that can adapt to future network fluctuations while meeting current image quality requirements. Based on this dynamically optimized target bitrate, the system further determines the appropriate target resolution and target frame rate for video encoding and transmission. This method effectively overcomes the poor adaptability and irrational resource allocation inherent in traditional solutions, which rely solely on instantaneous network status or ignore differences in video content. It aims to proactively ensure the smoothness and clarity of video conferencing in complex and changing network environments, enhancing the user experience.

[0091] Furthermore, a network-adaptive video conferencing transmission optimization system is also provided.

[0092] Figure 5 FIG. 1 is a block diagram of a network adaptive video conferencing transmission optimization system according to an embodiment of the present application. Figure 5As shown, the network adaptive video conference transmission optimization system 500 according to the embodiment of the present application includes: a network status data acquisition module 510, which always obtains the time set of network status historical data from the sending end; a current conference video frame type acquisition module 520, which is used to obtain the type of the current conference video frame; a network status data preprocessing module 530, which is used to preprocess the time set of the network status historical data to obtain a sequence of network status historical data; a video target bit rate generation module 540, which is used to determine the video target bit rate based on the type of the current conference video frame and the sequence of the network status historical data; a target resolution and frame rate determination module 550, which is used to determine the target resolution and target frame rate based on the video target bit rate and the type of the current conference video frame; a conference video frame encoding optimization module 560, which is used to encode the current conference video frame based on the target resolution and the target frame rate to obtain the adaptively optimized current conference video frame; and a conference video frame transmission module 570, which always transmits the adaptively optimized current conference video frame for data.

[0093] As described above, the network-adaptive video conferencing transmission optimization system 500 according to the embodiments of the present application can be implemented in various wireless terminals, such as a server equipped with a network-adaptive video conferencing transmission optimization algorithm. In one possible implementation, the network-adaptive video conferencing transmission optimization system 500 according to the embodiments of the present application can be integrated into the wireless terminal as a software module and / or hardware module. For example, the network-adaptive video conferencing transmission optimization system 500 can be a software module in the operating system of the wireless terminal, or an application developed specifically for the wireless terminal. Of course, the network-adaptive video conferencing transmission optimization system 500 can also be one of the many hardware modules of the wireless terminal.

[0094] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A network-adaptive video conferencing transmission optimization method, characterized in that: include: Obtain the time collection of historical network status data from the sender; Get the type of the current conference video frame; Preprocessing the time set of the network status historical data to obtain a sequence of network status historical data; Based on the type of the current conference video frame and the sequence of the network status historical data, determining the video target bit rate, including: arranging the sequence of the network status historical data into a network status historical data time series matrix according to the network status sample dimension and the time dimension; extracting the network status time series association feature from the network status historical data time series matrix to obtain a network status historical data time series association feature vector; performing one-hot encoding on the type of the current conference video frame to obtain a current conference video frame type one-hot encoding vector; passing the network status historical data time series association feature vector and the current conference video frame type one-hot encoding vector through a network status-video type progressive complementary interaction perception network to obtain a network status-video conference type progressive interaction feature vector; performing feature decoding on the network status-video conference type progressive interaction feature vector to obtain the video target bit rate; Determining a target resolution and a target frame rate based on the target video bit rate and the type of the current conference video frame; Encoding the current conference video frame based on the target resolution and the target frame rate to obtain an adaptively optimized current conference video frame; The adaptively optimized current conference video frame is transmitted for data transmission.

2. The network-adaptive video conferencing transmission optimization method according to claim 1, characterized in that: The network status historical data includes RTT mean, packet loss rate, receiving bit rate and sending bit rate.

3. The network-adaptive video conferencing transmission optimization method according to claim 2, characterized in that: The time set of the network status historical data is preprocessed to obtain a sequence of the network status historical data, including: normalizing and serializing the time set of the network status historical data to obtain the sequence of the network status historical data.

4. The network-adaptive video conferencing transmission optimization method according to claim 3, characterized in that: Extracting network state timing correlation features from the network state historical data timing matrix to obtain a network state historical data timing correlation feature vector, including: passing the network state historical data timing matrix through a network state timing correlation feature extractor based on a void convolutional neural network model to obtain the network state historical data timing correlation feature vector.

5. The network-adaptive video conferencing transmission optimization method according to claim 4, characterized in that: The network status historical data temporal correlation feature vector and the current conference video frame type one-hot encoding vector are passed through a network status-video type progressive complementary interaction perception network to obtain a network status-video conference type progressive interaction feature vector, including: Performing multi-level implicit feature extraction on the network status history data time series associated feature vector and the current conference video frame type one-hot encoding vector to obtain a network status history data time series middle-layer implicit feature encoding vector, a current conference video frame type middle-layer implicit feature encoding vector, a network status history data time series deep-layer implicit feature encoding vector, and a current conference video frame type deep-layer implicit feature encoding vector; Based on the middle-level implicit feature coding vector of the network status historical data time series, the middle-level implicit feature coding vector of the current conference video frame type, the deep-level implicit feature coding vector of the network status historical data time series and the deep-level implicit feature coding vector of the current conference video frame type, multi-level feature joint perception is performed on the network status historical data time series associated feature vector and the current conference video frame type one-hot coding vector to obtain the network status-video conference type progressive interaction feature vector.

6. The network-adaptive video conferencing transmission optimization method according to claim 5, characterized in that: Based on the middle-layer implicit feature coding vector of the network status history data time series, the middle-layer implicit feature coding vector of the current conference video frame type, the deep-layer implicit feature coding vector of the network status history data time series, and the deep-layer implicit feature coding vector of the current conference video frame type, multi-level feature joint perception is performed on the network status history data time series association feature vector and the current conference video frame type one-hot coding vector to obtain the network status-video conference type progressive interaction feature vector, including: Performing low-level feature fusion on the network status historical data time series correlation feature vector and the current conference video frame type one-hot encoding vector to obtain a network status-video conference type low-level fusion feature encoding vector; Performing high-level feature fusion on the implicit feature coding vector in the network status history data time series and the implicit feature coding vector in the current conference video frame type to obtain a network status-video conference type intermediate-level fusion feature coding vector and a network status-video conference type deep-level fusion feature coding vector; The network status-video conference type low-level fusion feature coding vector, the network status-video conference type middle-level fusion feature coding vector and the network status-video conference type deep-level fusion feature coding vector are progressively complementary and perceptually fused to obtain the network status-video conference type progressive interaction feature vector.

7. The network-adaptive video conferencing transmission optimization method according to claim 6, characterized in that: Performing high-level feature fusion on the implicit feature coding vector in the middle layer of the network status historical data time series and the implicit feature coding vector in the middle layer of the current conference video frame type to obtain a network status-video conference type middle-level fusion feature coding vector and a network status-video conference type deep-level fusion feature coding vector, including: Performing mid-level feature fusion on the mid-level implicit feature coding vector of the network status history data time series and the mid-level implicit feature coding vector of the current conference video frame type to obtain the mid-level fused feature coding vector of the network status-video conference type; Deep feature fusion is performed on the network status historical data time series deep implicit feature coding vector and the current conference video frame type deep implicit feature coding vector to obtain the network status-video conference type deep level fusion feature coding vector.

8. The network-adaptive video conferencing transmission optimization method according to claim 7, characterized in that: Feature decoding is performed on the network status-video conference type progressive interaction feature vector to obtain a video target bit rate, including: passing the network status-video conference type progressive interaction feature vector through a decoder-based target bit rate recommender to obtain a recommended decoding value of the video target bit rate.

9. Network adaptive video conferencing transmission optimization system, characterized by: include: The network status data acquisition module is used to obtain the time collection of network status historical data from the sending end; The current conference video frame type acquisition module is used to obtain the type of the current conference video frame; A network status data preprocessing module, configured to preprocess the time set of the network status historical data to obtain a sequence of network status historical data; The video target bit rate generation module always determines the video target bit rate based on the type of the current conference video frame and the sequence of the network status historical data, including: arranging the sequence of the network status historical data into a network status historical data time series matrix according to the network status sample dimension and the time dimension; extracting the network status time series association feature from the network status historical data time series matrix to obtain a network status historical data time series association feature vector; performing one-hot encoding on the type of the current conference video frame to obtain a current conference video frame type one-hot encoding vector; passing the network status historical data time series association feature vector and the current conference video frame type one-hot encoding vector through a network status-video type progressive complementary interaction perception network to obtain a network status-video conference type progressive interaction feature vector; and performing feature decoding on the network status-video conference type progressive interaction feature vector to obtain the video target bit rate. A target resolution and frame rate determination module, configured to determine a target resolution and a target frame rate based on the target video bit rate and the type of the current conference video frame; A conference video frame encoding optimization module for encoding the current conference video frame based on the target resolution and the target frame rate to obtain an adaptively optimized current conference video frame; The conference video frame transmission module is used to transmit the current conference video frame after the adaptive optimization.

Citation Information

Patent Citations

  • Video conference method, system and device and storage medium

    CN120050384A

  • Adaptive reward-driven video transport stream control method and device

    CN120075486A