Network adaptive video conference transmission optimization method and system
By analyzing the network status historical data and the current video frame type, dynamically optimizing the video target code rate, resolution and frame rate, the problems of inaccurate network status prediction and neglecting video content characteristics in the prior art are solved, and the fluency and clarity of video conferences are improved.
Patent Information
- Application Number
- CN202510830851.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The existing network adaptive video conferencing transmission optimization scheme has insufficient accuracy in network status prediction and cannot fully consider the characteristics of video content, resulting in poor adaptability and unreasonable resource allocation, and cannot maximize visual quality while ensuring fluency.
By analyzing the historical data sequence of network state, deep trends and patterns of network behavior are extracted, and combined with the type of current video frames, the fusion model is used for interactive perception and collaborative processing, and the video target code rate, resolution and frame rate are dynamically optimized to adapt to network fluctuations and meet picture quality requirements.
In a complex and changeable network environment, the forward-looking guarantee of the smoothness and clarity of video conferences is achieved, significantly improving the user experience.
Smart Images

Figure CN120358347A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network communication, and more specifically, to a method and system for optimizing video conference transmission with network adaptability. Background Art
[0002] With the rapid development of information technology and the increasing popularity of global collaboration, video conferencing systems have become an essential key tool in various scenarios such as enterprise communication, distance education, and online medical treatment. A high-quality and smooth video conferencing experience is crucial for ensuring communication efficiency and user satisfaction. However, the complexity and instability of Internet connections, such as bandwidth fluctuations, network congestion, packet loss, and transmission delays, often pose severe challenges to the transmission quality of video conferences, resulting in problems such as frozen frames, mosaics, and audio-video desynchronization, seriously affecting the user experience.
[0003] Currently, there are some network-adaptive video conference transmission optimization schemes. They usually adjust the bitrate, resolution, or frame rate of video encoding by monitoring network parameters (such as bandwidth, latency, packet loss rate). For example, some schemes may adopt rule-based control logic to reduce the bitrate when detecting network deterioration and then attempt to increase the bitrate when the network improves. Other schemes use traditional control theory models or simple statistical predictions to guide parameter adjustment. However, these existing schemes often have some deficiencies: one is that they mainly focus on the instantaneous performance or short-term changes of the network state, and insufficiently explore the historical trends and complex dynamic characteristics of the network state, resulting in inaccurate predictions and possible lags or oscillations in adaptive adjustments; the other is that in the decision-making process, they fail to fully consider the characteristics of the video content itself. For example, different types of video frames (such as I-frames, P-frames, B-frames, or more macroscopic content complexity such as static presentations and dynamic scenes) have different requirements for bitrate, resolution, and frame rate. Simply making a "one-size-fits-all" adjustment based on the network state cannot maximize the visual quality under specific content while ensuring smoothness, or waste unnecessary bandwidth when the content is simple.
[0004] Therefore, there is a need for a network-adaptive video conference transmission optimization scheme that can dynamically adjust transmission parameters according to the real-time network conditions to optimize the video conference transmission effect. Summary of the Invention
[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide a method and system for optimizing video conference transmission with network adaptability. By systematically analyzing the historical data sequence of the network state, deep trends and patterns of network behavior are extracted from it. At the same time, fully considering the different requirements of the current conference video frame type for transmission parameters (especially the bit rate), this key information is incorporated into the decision-making process. Through a specially designed fusion model, the historical evolution trend of the network state can be intelligently interactively perceived and collaboratively processed with the specific requirements of the current video content, so as to generate a video target bit rate that can not only adapt to future network fluctuations but also meet the current picture quality requirements. Based on this dynamically optimized target bit rate, the system further determines the appropriate target resolution and target frame rate, and encodes and transmits the video. This method can effectively overcome the problems of poor adaptability and unreasonable resource allocation caused by traditional solutions that simply rely on the instantaneous network state or ignore video content differences, aiming to proactively ensure the smoothness and clarity of video conferences and improve the user experience in a complex and changeable network environment.
[0006] According to one aspect of this application, a method for optimizing video conference transmission with network adaptability is provided, which includes:
[0007] Obtaining the time set of historical network state data from the sending end;
[0008] Obtaining the type of the current conference video frame;
[0009] Preprocessing the time set of the historical network state data to obtain a sequence of historical network state data;
[0010] Determining the video target bit rate based on the type of the current conference video frame and the sequence of the historical network state data;
[0011] Determining the target resolution and target frame rate based on the video target bit rate and the type of the current conference video frame;
[0012] Encoding the current conference video frame based on the target resolution and the target frame rate to obtain the adaptively optimized current conference video frame;
[0013] Transmitting the data of the adaptively optimized current conference video frame.
[0014] According to another aspect of this application, a system for optimizing video conference transmission with network adaptability is provided, which includes:
[0015] A network state data acquisition module, configured to obtain the time set of historical network state data from the sending end;
[0016] The current conference video frame type acquisition module is used to acquire the type of the current conference video frame;
[0017] The network status data preprocessing module is used to preprocess the time set of the historical network status data to obtain a sequence of historical network status data;
[0018] The video target bitrate generation module is always based on the type of the current conference video frame and the sequence of the historical network status data to determine the video target bitrate;
[0019] The target resolution and frame rate determination module is used to determine the target resolution and the target frame rate based on the video target bitrate and the type of the current conference video frame;
[0020] The conference video frame encoding optimization module is used to encode the current conference video frame based on the target resolution and the target frame rate to obtain the current conference video frame after adaptive optimization;
[0021] The conference video frame transmission module is used to perform data transmission on the current conference video frame after adaptive optimization.
[0022] Compared with the prior art, a network - adaptive video conference transmission optimization method and system provided by the present application systematically analyzes the historical data sequence of the network status, extracts the deep - level trends and patterns of network behavior from it. At the same time, fully considering the different requirements of the type of the current conference video frame for transmission parameters (especially the bitrate), this key information is incorporated into the decision - making process. Through a specially designed fusion model, it can intelligently interactively perceive and collaboratively process the historical evolution trend of the network status and the specific requirements of the current video content, so as to generate a video target bitrate that can not only adapt to future network fluctuations but also meet the current picture quality requirements. Based on this dynamically optimized target bitrate, the system further determines appropriate target resolution and target frame rate, encodes and transmits the video. This method can effectively overcome the problems of poor adaptability and unreasonable resource allocation caused by the traditional scheme's simple reliance on the instantaneous network status or ignoring video content differences, aiming to proactively ensure the smoothness and clarity of video conferences and improve the user experience in a complex and changeable network environment. Brief Description of the Drawings
[0023] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above - mentioned and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0024] Figure 1Flowchart of a network - adaptive video conferencing transmission optimization method according to an embodiment of the present application;
[0025] Figure 2 Schematic diagram of data flow of a network - adaptive video conferencing transmission optimization method according to an embodiment of the present application;
[0026] Figure 3 Flowchart of determining a video target bitrate based on the type of the current conference video frame and the sequence of the network state historical data in a network - adaptive video conferencing transmission optimization method according to an embodiment of the present application;
[0027] Figure 4 Flowchart of obtaining a network state - video conferencing type progressive interaction feature vector by passing the network state historical data temporal correlation feature vector and the one - hot encoding vector of the current conference video frame type through a network state - video type progressive complementary interaction perception network in a network - adaptive video conferencing transmission optimization method according to an embodiment of the present application;
[0028] Figure 5 Block diagram of a network - adaptive video conferencing transmission optimization system according to an embodiment of the present application. Detailed implementation manners
[0029] Next, example embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein.
[0030] As shown in the present application and the claims, unless the context clearly indicates an exception, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "including" and "comprising" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. A method or device may also include other steps or elements.
[0031] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server. The modules are only illustrative, and different aspects of the system and method can use different modules.
[0032] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations before or below do not necessarily need to be executed precisely in sequence. On the contrary, as needed, various steps can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several steps can be removed from these processes.
[0033] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0034] Aiming at the problems of insufficient accuracy in network state prediction and insufficient perception of video content characteristics in the existing solutions, this solution is no longer limited to a simple immediate response to the network state, but systematically analyzes the historical data sequence of the network state to extract the deep trends and patterns of network behavior. At the same time, fully considering the different requirements of the current conference video frame type for transmission parameters (especially the bit rate), this key information is incorporated into the decision-making process. Through a specially designed fusion model, it can intelligently interact and cooperate with the historical evolution trend of the network state and the specific requirements of the current video content, so as to generate a video target bit rate that can not only adapt to future network fluctuations but also meet the current picture quality requirements. Based on this dynamically optimized target bit rate, the system further determines the appropriate target resolution and target frame rate, and encodes and transmits the video. This method can effectively overcome the problems of poor adaptability and unreasonable resource allocation caused by the traditional solution's simple reliance on the instantaneous network state or ignoring video content differences, aiming to proactively ensure the smoothness and clarity of video conferences in a complex and changeable network environment and significantly improve the user experience.
[0035] In the technical solution of the present application, an optimized method for network-adaptive video conference transmission is proposed. Figure 1 It is a flowchart of an optimized method for network-adaptive video conference transmission according to an embodiment of the present application. Figure 2 It is a schematic diagram of data flow of an optimized method for network-adaptive video conference transmission according to an embodiment of the present application. As Figure 1 and Figure 2 shown, the optimized method for network-adaptive video conference transmission according to an embodiment of the present application includes the steps: S100, obtaining a time set of historical network state data from a sending end; S200, obtaining the type of the current conference video frame; S300, preprocessing the time set of the historical network state data to obtain a sequence of historical network state data; S400, determining a video target bit rate based on the type of the current conference video frame and the sequence of the historical network state data; S500, determining a target resolution and a target frame rate based on the video target bit rate and the type of the current conference video frame; S600, encoding the current conference video frame based on the target resolution and the target frame rate to obtain an adaptively optimized current conference video frame; S700, performing data transmission on the adaptively optimized current conference video frame.
[0036] Specifically, in steps S100 and S200, a time set of historical network status data is obtained from the sending end, and the type of the current conference video frame is obtained. The historical network status data includes the mean RTT, packet loss rate, received bitrate, and transmitted bitrate. It should be understood that in the network adaptive video conference transmission optimization, obtaining the time set of historical network status data from the sending end and obtaining the type of the current conference video frame are the basis and prerequisite for achieving precise dynamic adjustment. The reason for the need for these two aspects of information is that the historical data of the network status can reveal the dynamic characteristics and change trends of the network link, such as the long-term trends and periodic fluctuations of network behavior, providing a more reliable basis for predicting future network conditions, thereby making the bitrate decision more forward-looking. The type of the current conference video frame is directly related to the importance of the frame data and the characteristics of compression coding, reflecting the amount of information and compression requirements of the current picture. For example, key frames (I-frames) usually require a higher bitrate to ensure image quality, while predicted frames (P / B-frames) require a relatively lower bitrate. Judging based solely on the instantaneous network status is prone to interference from short-term factors such as network jitter, resulting in lag or overreaction in adaptive adjustment; without considering the type of video frame, it is impossible to optimize the transmission quality of key frames targeted or effectively save bandwidth when the content complexity is low. Therefore, combining the network historical trend with the current content characteristics can achieve the unity of network adaptability and content awareness, and dynamically allocate bitrate resources according to the content characteristics on the premise of ensuring network transmission stability, so as to continuously provide a better video conference experience in a complex and changing network environment.
[0037] More specifically, in a specific example of the present application, first, in terms of obtaining the time set of network status historical data at the sending end, its implementation generally relies on a continuous network parameter monitoring and recording mechanism. The sending end will periodically collect or obtain a series of key network metrics through feedback from the receiving end (such as through the RTCP protocol), such as the mean round-trip time (RTT), packet loss rate, actual receiving bitrate at the receiving end, and the sending bitrate of the sending end itself. These data points with timestamps will be recorded and stored in a buffer or database with a certain time window, forming a data set arranged in chronological order. For example, the system can set a sliding time window, such as the past few seconds or dozens of seconds, and continuously update the sequence of network status parameters within this window. When a decision needs to be made, relevant data is extracted from this historical data set. Secondly, in terms of obtaining the type of the current conference video frame, this is usually completed in the video encoding process or the pre-encoding processing stage. When the video encoder processes each video frame, it will classify it into different types according to its encoding strategy (such as GOP structure setting, scene change detection results, etc.), common types such as independently encoded key frames (I-frames), forward-predicted P-frames, or bidirectionally predicted B-frames. When a video frame is about to be processed for adaptive optimization decision-making, its type information can be directly obtained from the video encoder or video processing module. This information is crucial for determining the target bitrate of this frame subsequently, because different types of frames have very different requirements in terms of image quality assurance and compression ratio. Through the above methods, the sending end can simultaneously master the historical performance of the network link and the basic attributes of the current video content to be transmitted, laying a solid data foundation for subsequent intelligent determination of target bitrate, resolution, and frame rate based on this information.
[0038] Specifically, in step S300, the time set of the network status historical data is preprocessed to obtain a sequence of network status historical data. It should be understood that the originally collected network status historical data, such as the mean RTT, packet loss rate, receiving bitrate, and sending bitrate, often have different physical units, dimensions, and numerical ranges. For example, the RTT may be in milliseconds and fluctuate between dozens and hundreds, while the packet loss rate is a percentage and the value is between 0 and 100 (or 0 and 1). If these heterogeneous data are directly input into the subsequent complex model without processing, the features with larger numerical ranges may disproportionately dominate the learning process of the model, masking the contributions of other features that are numerically smaller but equally important, resulting in difficult model training, slow convergence speed, and even inability to accurately capture the true relationships between features. In addition, these data points occur in time, and their inherent time dependence is crucial for predicting future network trends. The original "time set of network status historical data" needs to be transformed into a structured "sequence" that can clearly reflect this temporal relationship in order to be effectively utilized by the time series model.
[0039] Specifically, in the embodiments of the present application, the time set of the network state historical data is preprocessed to obtain a sequence of network state historical data, including: performing normalization processing and serialization processing on the time set of the network state historical data to obtain the sequence of network state historical data. That is to say, first, through normalization processing, network state data in different dimensions (RTT, packet loss rate, etc.) can be uniformly mapped to a similar numerical interval, such as the common [0, 1] or [-1, 1] interval. This can eliminate the influence of dimensionality and the difference in value ranges between different features, enabling the model to treat each input feature fairly, avoiding learning biases caused by numerical size differences, thereby accelerating the convergence speed of model training and enhancing the stability and generalization ability of the model. Second, through serialization processing, a set of data points that may originally only carry timestamps is strictly organized in chronological order to form an ordered time series. This provides a structured input for extracting temporal correlation features of network states subsequently, ensuring that the model can effectively learn and capture key information such as the dynamic pattern, periodic law, and mutation of network states evolving over time.
[0040] Specifically, in step S400, based on the type of the current conference video frame and the sequence of the network state historical data, the video target bitrate is determined. It should be understood that traditional bitrate control strategies often focus on the response to the current instantaneous network state or fail to fully distinguish the characteristics of different video contents, which leads to problems such as insufficient adaptability, inaccurate prediction, and suboptimal resource allocation. Moreover, the sequence of network state historical data contains deep information such as the trend, periodicity, and suddenness of network behavior, which can reveal the true carrying capacity of the network link and possible future fluctuations far better than a single instantaneous data point. At the same time, there are significant differences in the bitrate requirements for different types of current conference video frames, such as key frames (I frames) with large amounts of information and high quality requirements and predictive frames (P frames or B frames) with relatively small amounts of information and strong compressibility. Therefore, it is necessary to effectively combine these two dimensions, namely the long-term dynamic characteristics of the network and the immediate encoding requirements of video content, to break the limitations of traditional solutions. In particular, through in-depth analysis of the network historical state and keen perception of the characteristics of the current video content, an "optimal" video target bitrate that can adapt to the current and foreseeable future network conditions and meet the quality requirements of specific video frames is dynamically calculated. This "optimal" does not simply pursue the highest bitrate but rather allocates reasonable bitrate resources for video frames with different importance levels and compression characteristics on the premise of ensuring transmission fluency (i.e., the network can carry), thereby seeking the best balance between video clarity and fluency in a changing network environment. This precisely determined target bitrate will serve as the key basis for subsequent adjustment of video resolution and frame rate and ultimately guide the video encoding process.
[0041] Figure 3 It is a flowchart for determining a target resolution and a target frame rate based on the video target bitrate and the type of the current conference video frame of the network-adaptive video conference transmission optimization method according to an embodiment of the present application. As Figure 3 shown, for the network-adaptive video conference transmission optimization method according to an embodiment of the present application, step S400 includes: S410, organizing the sequence of the network state historical data into a network state historical data time series matrix according to the network state sample dimension and the time dimension; S420, extracting network state time series correlation features from the network state historical data time series matrix to obtain a network state historical data time series correlation feature vector; S430, performing one-hot encoding on the type of the current conference video frame to obtain a current conference video frame type one-hot encoding vector; S440, passing the network state historical data time series correlation feature vector and the current conference video frame type one-hot encoding vector through a network state-video type progressive complementary interaction perception network to obtain a network state-video conference type progressive interaction feature vector; S450, performing feature decoding on the network state-video conference type progressive interaction feature vector to obtain a video target bitrate.
[0042] Specifically, in step S410, the sequence of the network state historical data is organized into a network state historical data time series matrix according to the network state sample dimension and the time dimension. It should be understood that although the "sequence of network state historical data" obtained by the previous preprocessing has solved the basic problems of data heterogeneity and time series, for many advanced analysis models, especially those deep learning models designed to capture complex time dependencies and multivariate interactions, directly using a one-dimensional sequence or a simple list structure may not be efficient enough or cannot fully utilize the inherent structure of the data. The network state itself is multi-dimensional ("network state sample dimension", such as multiple metrics like RTT, packet loss rate, sending bit rate, receiving bit rate, etc.), and these multi-dimensional metrics evolve over time ("time dimension"). Therefore, in the technical solution of this application, the sequence of the network state historical data is further organized into a network state historical data time series matrix according to the network state sample dimension and the time dimension. It is worth mentioning that organizing these multi-dimensional network state historical time series data into a matrix, where one axis represents the time step and the other axis represents different network state parameters, can present the historical dynamics of the network in a highly structured manner. This network state historical data time series matrix can clearly show the specific values of multiple network parameters at consecutive time points and their mutual relationships. This matrix form not only facilitates subsequent mathematical operations and feature transformations, but also can completely retain and explicitly express the distribution information of the network state on the two-dimensional plane of "feature - time". This provides an ideal data basis for further exploring the potential, non-linear time series correlation features in the network state, such as periodic fluctuations, trend changes, burst congestion patterns, etc.
[0043] Specifically, in step S420, network state time-series correlation features are extracted from the network state historical data time-series matrix to obtain a network state historical data time-series correlation feature vector. It should be understood that although the original network state historical data time-series matrix structurally presents the changes of multi-dimensional network parameters over time, it is still a relatively raw data representation. Directly using the entire high-dimensional time-series matrix for subsequent decision fusion not only incurs high computational costs but also may contain redundant information or noise. More importantly, it fails to explicitly reveal the deep-seated, non-linear correlation patterns and long-term dependencies between different time points and different network parameters, which are the keys to accurately judging network trends and predicting future states. Traditional statistical methods or simple time-series analysis are difficult to fully exploit these complex internal connections. Therefore, in the technical solution of this application, network state time-series correlation features are extracted from the network state historical data time-series matrix to obtain a network state historical data time-series correlation feature vector. The high-dimensional network state historical data time-series matrix is mapped and condensed into a low-dimensional but higher-information-content network state historical data time-series correlation feature vector. This feature vector aims to encapsulate and characterize the most critical and representative time-series dynamic characteristics in network historical data, such as the evolution pattern of network congestion, the periodic law of bandwidth fluctuations, and the potential causal or concomitant relationships between different parameters (such as RTT and packet loss rate). Dilated convolutional neural networks are very suitable for learning these complex correlation features from network time-series data because they can effectively capture long-term dependencies in time series by adjusting the dilation rate without adding excessive computational amounts and parameters. Therefore, in the embodiment of this application, extracting network state time-series correlation features from the network state historical data time-series matrix to obtain a network state historical data time-series correlation feature vector includes: passing the network state historical data time-series matrix through a network state time-series correlation feature extractor based on a dilated convolutional neural network model to obtain the network state historical data time-series correlation feature vector. Its goal is to generate a compact feature representation that can accurately reflect the overall historical situation of the network and highlight key change trends, providing a high-quality, refined network state "portrait" for subsequent intelligent fusion with video frame type information. This network state historical data time-series correlation feature vector can capture the essence of network dynamics more effectively than the original time-series data or simple statistical features. This feature vector is no longer just a list of data but an abstraction and generalization of network behavior patterns, eliminating noise and redundancy and strengthening key information. This enables more accurate interaction based on the network historical trend during the subsequent fusion with video frame type features in the network state-video type progressive complementary interaction perception process, thereby generating a more reasonable and forward-looking video target bitrate.
[0044] Specifically, in step S430, the type of the current conference video frame is one-hot encoded to obtain a one-hot encoded vector of the current conference video frame type. It should be understood that subsequent intelligent decision-making models usually require numerical inputs rather than the original category labels (such as "I-frame", "P-frame", "B-frame"). If these category labels are simply replaced with arbitrary numbers (for example, I = 1, P = 2, B = 3), an ordinal relationship or magnitude comparison that does not exist will be inadvertently introduced, misleading the model into thinking that the P-frame is "greater" than the I-frame in a certain sense, which does not conform to the actual characteristics of the video frame type and will interfere with the learning process and decision-making accuracy of the model. One-hot encoding can transform these discrete and unordered category features into a numerical and high-dimensional sparse vector form that is easy for the model to understand and process, avoiding this artificially introduced bias. Therefore, in the technical solution of this application, the type of the current conference video frame is further one-hot encoded to obtain a one-hot encoded vector of the current conference video frame type. In this way, an important content feature, that is, the type of the current conference video frame, can be characterized in a way that is friendly and unbiased to the machine learning model. Through one-hot encoding, each video frame type (such as I-frame, P-frame, B-frame) will be mapped to a new vector space, and the dimension of the vector is equal to the total number of frame types. The dimension corresponding to the current frame type is 1, and the remaining dimensions are all 0. For example, if there are three frame types, the I-frame may be encoded as [1, 0, 0], the P-frame as [0, 1, 0], and the B-frame as [0, 0, 1]. This ensures that different frame types are independent and equal category identifiers when input into the model, without any numerical magnitude or order implication, accurately reflecting their essential differences as different content units. This one-hot encoded vector will subsequently be used as an input together with the feature vector extracted from the network state historical data for subsequent fusion analysis by the interaction perception network.
[0045] Specifically, in step S440, the network state historical data temporal correlation feature vector and the current conference video frame type one-hot encoding vector are passed through the network state-video type progressive complementary interaction perception network to obtain the network state-video conference type progressive interaction feature vector. It should be understood that simply concatenating or linearly combining the extracted network state historical temporal correlation features and the current conference video frame type features is far from sufficient to capture the profound and non-linear dependence relationship between the two, and this relationship is crucial for optimizing video conference transmission. The dynamic changes in the network environment (characterized by the network state historical data temporal correlation feature vector) have different impacts on the transmission quality of different types of video frames (characterized by the current conference video frame type one-hot encoding vector). Conversely, different video frames also have different demands for network resources. For example, a stable high-bandwidth network environment is extremely beneficial for key I-frames, and a high bitrate can be allocated to ensure clarity; while for a network with large fluctuations, even if the current instantaneous state is acceptable, a more conservative bitrate strategy may need to be adopted for upcoming P-frames or B-frames. This complex, multi-level interaction judgment cannot be effectively modeled by a simple model. Therefore, in the technical solution of this application, the network state historical data temporal correlation feature vector and the current conference video frame type one-hot encoding vector are further passed through the network state-video type progressive complementary interaction perception network to obtain the network state-video conference type progressive interaction feature vector. Through the processing of the network state-video type progressive complementary interaction perception network, not only are the network state and video frame type regarded as independent factors, but also the synergistic effects and complementary information at different abstraction levels are deeply mined. This network, through its multi-level, phased processing paradigm - first independently performs multi-level hidden feature extraction on the network state historical data temporal correlation feature vector and the current conference video frame type one-hot encoding vector (generating their respective low-level, middle-level, and deep-level hidden features), then performs feature fusion at the corresponding abstraction levels (the low level focuses on detail correspondence, the middle level focuses on structural association, and the deep level focuses on semantic synergy), and finally integrates the interaction patterns captured at these different levels through a progressive complementary perception fusion mechanism. This final "network state-video conference type progressive interaction feature vector" aims to generate a more comprehensive and robust joint representation, which comprehensively reflects the deep coupling relationship between the current network historical trend and the current video content demand, providing a high-quality, high-information-density decision basis for the determination of the downstream video target bitrate. Through this progressive complementary interaction perception, the system can transcend the simple response to a single factor and achieve the combination of "content awareness" and "network prediction" in a true sense.For example, the network can understand that even if the deep features of the network state history indicate that the network tends to be stable (the result of deep fusion), but if the current frame is a P-frame (the result of low / middle layer feature interaction) and there is slight jitter recently (the result of middle layer feature interaction), then the final interactive feature vector will guide a compromise bitrate that takes advantage of both the network stability dividend and the characteristics of the P-frame and recent small fluctuations. This effectively avoids problems such as inaccurate adjustment and insufficient resource utilization caused by ignoring the network historical trend and the influence of video frame types, and ultimately maximally guarantees the clarity and smoothness of video conferencing in a complex and changeable network environment.
[0046] Figure 4 The flowchart for obtaining the video stream segment-audio semantic search response coding vector by passing the sequence of the audio mel spectrogram semantic feature vector and the video stream segment semantic feature vector through the graph learning-based video segment semantic feature search module for the network adaptive video conferencing transmission optimization method according to an embodiment of the present application. As Figure 4 shown, for the network adaptive video conferencing transmission optimization method according to an embodiment of the present application, step S440 includes: S441, performing multi-level hidden feature extraction on the network state history data temporal correlation feature vector and the current conference video frame type one-hot encoding vector to obtain a network state history data temporal middle layer hidden feature encoding vector, a current conference video frame type middle layer hidden feature encoding vector, a network state history data temporal deep layer hidden feature encoding vector, and a current conference video frame type deep layer hidden feature encoding vector; S442, based on the network state history data temporal middle layer hidden feature encoding vector, the current conference video frame type middle layer hidden feature encoding vector, the network state history data temporal deep layer hidden feature encoding vector, and the current conference video frame type deep layer hidden feature encoding vector, performing multi-level feature joint perception on the network state history data temporal correlation feature vector and the current conference video frame type one-hot encoding vector to obtain the network state-video conference type progressive interaction feature vector.
[0047] More specifically, in step S441, performing multi-level hidden feature extraction on the network state history data temporal correlation feature vector and the current conference video frame type one-hot encoding vector to obtain a network state history data temporal middle layer hidden feature encoding vector, a current conference video frame type middle layer hidden feature encoding vector, a network state history data temporal deep layer hidden feature encoding vector, and a current conference video frame type deep layer hidden feature encoding vector, which is expressed by the formula as:
[0048]
[0049]
[0050]
[0051]
[0052] Among them, is the time-series correlation feature vector of the network state historical data, and are respectively the trainable middle-layer weight matrix of the network state historical data time series and the trainable middle-layer bias vector of the network state historical data time series, is the middle-layer hidden feature encoding vector of the network state historical data time series, is the one-hot encoding vector of the current conference video frame type, and are respectively the trainable middle-layer weight matrix of the current conference video frame type and the trainable middle-layer bias vector of the current conference video frame type, is the middle-layer hidden feature encoding vector of the current conference video frame type, and are respectively the trainable deep-layer weight matrix of the network state historical data time series and the trainable deep-layer bias vector of the network state historical data time series, is the deep-layer hidden feature encoding vector of the network state historical data time series, and are respectively the trainable deep-layer weight matrix of the current conference video frame type and the trainable deep-layer bias vector of the current conference video frame type, is the deep-layer hidden feature encoding vector of the current conference video frame type.
[0053] It should be understood that directly using the time-series correlation feature vector of the original network state history data and the one-hot encoding vector of the current conference video frame type for single-level interaction cannot fully capture the complex correlations between the two at different levels of abstraction. The time-series features of the network state and the content attributes of the video frames do not interact in a single dimension. For example, the correlation between the long-term stability of the network and the global importance of the frame type (such as I-frame), and the correlation between the short-term burst jitter of the network and the local encoding requirements of the frame type (such as P-frame) belong to different levels of interaction. Therefore, it is necessary to first perform independent deep non-linear transformations on these two input feature vectors respectively, project them into different feature subspaces, so as to reveal their inherent and deeper characteristics at multiple levels of abstraction and prepare for subsequent targeted interactions. Through the hierarchical feature learning ability of the deep neural network, the time-series correlation feature vector of the network state history data is transformed to obtain the "time-series middle-level hidden feature encoding vector of the network state history data" that can represent the medium-term trend and structural fluctuations of the network, and the "time-series deep-level hidden feature encoding vector of the network state history data" that can represent the long-term situation and global semantics of the network. Similarly, the one-hot encoding vector of the current conference video frame type is transformed to obtain the "middle-level hidden feature encoding vector of the current conference video frame type" that can reflect the typical requirements of this frame type in medium-complexity scenarios (such as in a GOP structure), and the "deep-level hidden feature encoding vector of the current conference video frame type" that reflects its contribution to the overall video quality and fundamental encoding requirements. These hidden feature vectors at different levels respectively encapsulate the core information of the original input at a specific abstraction granularity, laying a solid foundation for more accurate and targeted feature interaction at the corresponding levels in the future.
[0054] Correspondingly, according to the embodiment of the present application, step S442 includes: performing low-level feature fusion on the time-series correlation feature vector of the network state history data and the one-hot encoding vector of the current conference video frame type to obtain a low-level fusion feature encoding vector of the network state-video conference type; performing high-level feature fusion on the time-series middle-level hidden feature encoding vector of the network state history data and the middle-level hidden feature encoding vector of the current conference video frame type to obtain a middle-level fusion feature encoding vector and a deep-level fusion feature encoding vector of the network state-video conference type; performing progressive complementary perception fusion on the low-level fusion feature encoding vector of the network state-video conference type, the middle-level fusion feature encoding vector of the network state-video conference type, and the deep-level fusion feature encoding vector of the network state-video conference type to obtain the progressive interaction feature vector of the network state-video conference type.
[0055] More specifically, perform low-level feature fusion on the time-series correlation feature vector of the network state historical data and the one-hot encoding vector of the current conference video frame type to obtain a low-level fusion feature encoding vector of network state-video conference type, which is expressed by the formula:
[0056]
[0057] where is element-wise addition, is a multi-layer perceptron, is the low-level fusion feature encoding vector of network state-video conference type.
[0058] It should be understood that although there will be middle-level and deep-level feature extraction and fusion subsequently, these deep abstraction processes will inevitably smooth or lose some high-frequency details and the most direct correspondence relationships in the original input features. Although the time-series correlation feature vector of the network state historical data has been refined once, it still contains relatively direct network fluctuation details inside, while the one-hot encoding vector of the current conference video frame type directly indicates the fundamental attributes of the current frame. To ensure that the final decision can make full use of these most primitive and direct interaction information and prevent the loss of sensitivity to instantaneous changes or precise matches during subsequent abstraction processes, it is necessary to perform an "early fusion" at the lowest level, that is, as close as possible to the level of the input features. In the scenario of video conferencing, through low-level feature fusion, this means directly capturing the most immediate and subtle fluctuations in the network state history (such as instantaneous spikes in RTT, bursts of packet loss) and the direct, less abstracted correlation between the specific type of the current frame being an I-frame, P-frame, or B-frame. For example, the interaction between a just-occurred network congestion indication (high-frequency information) and the current being an I-frame sensitive to packet loss (precise correspondence) can be most directly reflected in the low-level fusion. This low-level fusion feature encoding vector aims to encapsulate this most primitive and basic correspondence pattern, providing a perspective that retains the original details for subsequent progressive complementary perception fusion.
[0059] Accordingly, according to the embodiments of the present application, high-level feature fusion is performed on the middle-level implicit feature encoding vectors in the network state historical data time series and the middle-level implicit feature encoding vectors in the current conference video frame type to obtain the middle-level fusion feature encoding vectors of the network state-video conference type and the deep-level fusion feature encoding vectors of the network state-video conference type, including: performing middle-level feature fusion on the middle-level implicit feature encoding vectors in the network state historical data time series and the middle-level implicit feature encoding vectors in the current conference video frame type to obtain the middle-level fusion feature encoding vectors of the network state-video conference type; performing deep feature fusion on the deep-level implicit feature encoding vectors in the network state historical data time series and the deep-level implicit feature encoding vectors in the current conference video frame type to obtain the deep-level fusion feature encoding vectors of the network state-video conference type.
[0060] More specifically, performing middle-level feature fusion on the middle-level implicit feature encoding vectors in the network state historical data time series and the middle-level implicit feature encoding vectors in the current conference video frame type to obtain the middle-level fusion feature encoding vectors of the network state-video conference type, which is expressed by the formula as:
[0061] ;
[0062]
[0063] Wherein, is the attention fusion process, is the element-wise multiplication by position, is the scale of, is function, is the middle-level fusion feature encoding vector of the network state-video conference type.
[0064] It should be understood that although low-level fusion retains the original details, it may not be enough to capture more macroscopic and structural association patterns, while jumping directly to deep semantic fusion may ignore those important intermediate-scale interactions between details and the global. The mutual influence between the historical evolution of network state (such as trends and periodic fluctuations within a few seconds to tens of seconds) and video frame type (considering its position in the GOP structure, dependence on previous and subsequent frames and other mid-level attributes) is often most significant at a "structural" level that goes beyond instantaneous correspondence but does not reach global semantics. Among them, the mid-level features have filtered out some noise and encoded more stable patterns, so fusion at this level can more effectively identify and utilize the interactions between these initially abstracted and refined stable patterns. In particular, here, the mid-level implicit features of network state represent structural information such as the formation and relief patterns of network congestion and the periodic fluctuations of bandwidth; while the mid-level implicit features of video frame type reflect its role in the coding unit (such as GOP), its degree of dependence on reference frames, or its mid- and short-term contribution to overall visual coherence. The purpose of mid-level fusion is to model the structural association or composition relationship across sources and capture this more complex dependency than the low-level. For example, it can learn how to adjust the target bitrate for a P frame (mid-level feature of the video frame, reflecting its structural dependency) that follows a key I frame when the network shows a slight deterioration trend that lasts for several seconds (mid-level feature of the network). This judgment is more robust than simply responding to instantaneous network spikes (low-level) and more sensitive than only considering the long-term average network quality (deep level).
[0065] More specifically, the network status historical data temporal deep implicit feature coding vector and the current conference video frame type deep implicit feature coding vector are deeply fused to obtain the network status-video conference type deep level fusion feature coding vector, which is expressed as:
[0066]
[0067]
[0068] in, is the trainable deep weight matrix, for Activation function, and is the low-rank projection matrix, for Activation function, The network status-video conference type deep level fusion implicit feature encoding vector, Deep level fusion feature encoding vector for network status-video conference type.
[0069] It should be understood that merely retaining low-level details and capturing mid-level structural associations is not sufficient to form a global and strategic understanding of the network state and video content requirements. The long-term trends of the network state, overall health (e.g., average bandwidth, packet loss rate, stability assessment over several minutes or longer), and the fundamental importance of video frame types in the entire video stream or their core contribution to the ultimate interactive perception quality (e.g., the absolute criticality of I-frames as the decoding starting point, or the sacrificeability of B-frames in specific scenarios) all belong to the "deep semantics" that can only be grasped from a more macroscopic and abstract level. Therefore, it is necessary to fuse the deep features of the network state, which have been maximally abstracted and information-compressed, with the deep features of video frame types in order to model the synergistic effect between the two at the highest level. In a video conferencing scenario, it is necessary to understand how to ensure the basic transmission quality of key I-frames (deep features of video frames, representing their core semantic value) in the case of a long-term poor overall network environment (deep features of the network), or how to maximize the encoding quality of all types of frames to improve the overall visual experience when the network is long-term stable and of high quality. This deep fusion aims to extract a highly generalized and guiding interactive judgment, which pays more attention to long-term strategies rather than short-term tactics, and sets a macroscopic tone or constraint for the entire bitrate adaptation system.
[0070] More specifically, according to an embodiment of the present application, progressive complementary perceptual fusion is performed on the low-level fusion feature encoding vector of the network state - video conferencing type, the mid-level fusion feature encoding vector of the network state - video conferencing type, and the deep-level fusion feature encoding vector of the network state - video conferencing type to obtain the progressive interactive feature vector of the network state - video conferencing type, which is expressed by the formula:
[0071]
[0072]
[0073] Wherein, is and cascading operation of, is a trainable cascading weight matrix, is function, is the shallow - mid-level fusion gating adjustment parameter of the network state - video conferencing type, is the shallow - mid-level fusion feature encoding vector of the network state - video conferencing type, 、 and are a trainable query matrix, a trainable key matrix, and a trainable value matrix respectively, 、 and They are the shallow-medium level fusion query feature vector of network status-video conference type, the deep-level fusion key feature vector of network status-video conference type, and the medium-level fusion value feature vector of network status-video conference type. for The scale of is the network status-video conference type progressive interaction feature vector.
[0074] It should be understood that the fusion feature encoding vectors of network status-video conference type generated at the three levels of low, medium and deep respectively capture the interactive information of network status and video frame type from different granularities (instantaneous details, structural patterns, global semantics), but they are independent and only interpret this complex relationship from a single perspective. If they are simply spliced or averaged, it may lead to information redundancy, dilution of important features, or inability to effectively resolve potential conflicts between different levels, and even "performance degradation" may occur. Therefore, it is necessary to use the method of progressive complementary perceptual fusion to intelligently integrate these multi-level interactive information to ensure that they can enhance each other rather than interfere with each other. In this way, a final unified representation with a comprehensive, profound and collaborative understanding of the interactive relationship between network status and video content can be constructed, namely, the "progressive interactive feature vector of network status-video conference type". The "progressive" here means that the fusion process is orderly and progressive, and the low-level information will provide a basis for the middle-level information, and the middle-level information will provide context for the high-level information, that is, it is gradually integrated through a dynamic weighting or gating mechanism. "Complementarity" emphasizes the uniqueness and irreplaceability of features at different levels. They each contribute different aspects of the overall understanding. The fusion goal is to complement each other and form a more powerful and comprehensive feature representation than any single level. This fusion is committed to dynamically integrating interactive information at different levels of abstraction, so that the final network state-video conference type progressive interaction feature vector can take into account the instantaneous changes, medium-term trends and long-term trends between network status and video conference type, and cleverly combine it with the immediate needs, structural importance and core value of video frames. In the video conferencing scenario, this means that the system can make extremely accurate and robust judgments based on this feature vector: for example, when the low level shows a short network spike, but the middle and deep levels indicate that the network is stable as a whole and the current frame is a critical I frame, the progressive complementary fusion can weigh this information and decide to moderately smooth the impact of the spike, give priority to ensuring the quality of the I frame, rather than blindly and drastically reduce the code. On the contrary, if multiple levels of information all point to network deterioration, a more conservative bit rate strategy will be decisively adopted. This intelligent integration capability enables the system to demonstrate excellent adaptability, foresight and stability when facing a complex, changeable and uncertain real network environment, ultimately significantly improving the fluency and visual quality of video conferencing, ensuring users a high-quality experience under various network conditions.
[0075] Preferably, when performing step-by-step complementary perception aggregation on the low-level fusion feature encoding vector of the network state - video conference type, the middle-level fusion feature encoding vector of the network state - video conference type, and the deep-level fusion feature encoding vector of the network state - video conference type, the low-level fusion feature encoding vector of the network state - video conference type and the middle-level fusion feature encoding vector of the network state - video conference type substantially generate the shallow-middle level fusion gating adjustment parameters of the network state - video conference type , and then perform fusion based on the phase complementarity of the shallow-middle level fusion gating adjustment parameters of the network state - video conference type , and at the same time further interact with the deep-level fusion feature encoding vector of the network state - video conference type via the weight matrices , and to perform the restriction of the order parameter field to promote the interaction. Here, due to the instability of the solution of the binding structure of the order parameter field, the interaction is hindered, thus affecting the representation performance of the progressive interaction feature vector of the network state - video conference type.
[0076] Therefore, for the pre-trained matrices , and , first calculate the path integral representation of the decomposition field of the order parameter field jointly formed by each matrix:
[0077]
[0078]
[0079]
[0080] wherein, , and are all integral operations, , and are all , and 's path integral representation matrices of the order parameter field decomposition field, and the above multiplications are all matrix dot multiplications.
[0081] Then, define the field uniform state alignment loss function :
[0082]
[0083] where denotes the nuclear norm of the matrix, i.e., the sum of the eigenvalues of the matrix, is the scaling hyperparameter.
[0084] In this way, through the path integral of each decomposed field of the order parameter field, the non-integrable superposition of the bound connections of the order parameter field is weakened to prevent the geometric phase superposition that hinders the formation of the interaction of the local field structure. Then, the instability of the bound structure solution can be avoided through the field uniform state equivalence of the sum of eigenvalues, thereby realizing the , and maintenance of the interaction, and improving the characterization performance of the eigenvector of the progressive interaction of the network state-video conference type.
[0085] Specifically, in step S500, based on the video target bitrate and the type of the current conference video frame, the target resolution and the target frame rate are determined. It should be understood that the clarity (mainly determined by the resolution) and the smoothness (mainly determined by the frame rate) of the video are mutually restrictive and both consume bitrate. When the total bitrate is limited, a trade-off must be made between the two. Different types of video frames (such as I frames, P frames, B frames) have significantly different characteristics for the final decoding quality and bitrate consumption. For example, as a key frame, the quality of the I frame has a significant impact on subsequent frames and usually requires a higher bitrate to ensure its clarity, while the B frame has a higher fault tolerance. Therefore, dynamically adjusting the resolution and the frame rate according to the target bitrate and the current frame type aims to most effectively allocate the limited bitrate resources to the dimension where the current frame most needs them, in order to maximize the user's visual experience under the premise of meeting the transmission constraints.
[0086] More specifically, in a specific example of the present application, in the first step, a bitrate-driven preliminary parameter range is defined. The system will, according to the input target video bitrate, refer to a pre-set bitrate-resolution-frame rate correspondence table or empirical model to preliminarily screen out a group or a roughly suitable resolution and frame rate range. For example, an extremely low target bitrate (such as a few hundred kbps) usually corresponds to a lower resolution (such as 360p or even lower) and possibly a lower frame rate (such as 15-20fps), while a higher target bitrate (such as several Mbps) allows a higher resolution (such as 720p or 1080p) and a standard frame rate (such as 25-30fps). This step mainly ensures that the selected parameters can be theoretically carried by the target bitrate. In the second step, parameter bias adjustment is performed according to the type of the current conference video frame. At this stage, the system will refine the selection. If the current frame is an I-frame, considering its key role as a decoding reference, the system tends to preferentially ensure a higher resolution within the bitrate budget to ensure the basic clarity of the picture. At this time, if the bitrate is very tight, the frame rate will be slightly sacrificed (but the minimum acceptable smoothness needs to be maintained). If the current frame is a P-frame, the system will seek a relatively balanced distribution between the resolution and the frame rate because the P-frame carries most of the motion information and has certain requirements for smoothness. If the current frame is a B-frame, due to its non-reference and discardable nature, when the bitrate is tight, the system will more significantly reduce its resolution or even the frame rate to save bitrate for more important frames or to provide buffering during network fluctuations. This adjustment may also consider content characteristics. For example, for high-dynamic scenes, the priority of the frame rate may be increased; for conferences with more static content, the priority of the resolution may be higher. In the third step, final decision-making and constraint checking are performed. Based on the analysis of the first two steps, the system will select a specific combination of target resolution and target frame rate. This decision may be made by referring to a more refined predefined strategy table (for example, for different bitrate levels and different frame type combinations, the optimal resolution / frame rate pairs are preset), or by evaluating the expected perceived quality of different combinations under the current conditions through a simple utility function. At the same time, feasibility checks will also be carried out to ensure that the selected parameter combination meets the capabilities of the encoder and certain minimum quality standards (for example, the frame rate is not lower than a certain threshold to avoid picture stuttering, and the resolution is not lower than a certain threshold to ensure basic recognizability).
[0087] Specifically, in step S600, the current conference video frame is encoded based on the target resolution and the target frame rate to obtain the current conference video frame after adaptive optimization. It should be understood that the resolution determines the clarity of the image, and the frame rate determines the smoothness of the motion. These two parameters are the key factors for resource consumption during video encoding. Setting them in advance means that when the encoder starts processing each frame, it already knows the spatial dimension of the output image and the number of frames per unit time, so that it can adopt encoding strategies and allocate bits accordingly to ensure that the generated video stream not only meets the transmission capacity but also satisfies the user's core requirements for clarity and fluency. Therefore, in the technical solution of the present application, the current conference video frame is encoded based on the target resolution and the target frame rate to obtain the current conference video frame after adaptive optimization. This can effectively convert the limited channel resources (characterized by the previously determined "video target bit rate") into video quality that can be perceived by the user, that is, under dynamically changing network conditions, by adjusting the spatial details and temporal fluency of the video stream, strive to achieve the best balance between the encoding and transmission costs and the final visual experience.
[0088] More specifically, in a specific example of the present application, the first step is frame capture and parameter adaptation preprocessing. The system obtains the original video frames from a video capture device (such as a camera). These original frames usually have the resolution and frame rate inherent to the device. Subsequently, these original frames must be adjusted to match the previously determined "target resolution" and "target frame rate". If the original resolution is higher than the target resolution, image downsampling (such as using bilinear interpolation, bicubic interpolation, or more advanced learning-based scaling algorithms) is required to reduce each frame image to the target size; conversely, if the target resolution is higher (less common in bitrate-limited adaptive scenarios, unless there are specific super-resolution requirements), upsampling may be needed. Similarly, for the frame rate, if the original frame rate is higher than the target frame rate, temporal frame skipping is required, for example, simply discarding the redundant frames or adopting a more intelligent frame selection strategy to ensure the coherence of motion information; if the target frame rate is lower than the original frame rate, no special processing is required, and only the frames need to be sent at the target rate (in real-time coding, generally, frame interpolation to increase the frame rate is not involved because of its large computational amount and possible introduction of unnaturalness). This step ensures that the data fed into the core encoder conforms to the instructions of "adaptive optimization" in both spatial and temporal dimensions. The second step is core video coding compression. The preprocessed video frames with the target resolution and target frame rate are fed into the video encoder. The encoder performs a series of complex operations according to the selected coding standard to remove redundant information, mainly including: performing intra-frame prediction using the spatial correlation of the pixels within the frame, or performing inter-frame prediction using the temporal correlation between consecutive frames (this process will use the type of the current frame, such as I, P, B frames, to guide the selection of the prediction mode); transforming the prediction residuals (such as DCT) to concentrate the energy; quantizing the transform coefficients, which is the key to lossy compression and is directly constrained by the target bitrate (dynamically adjusting the quantization parameter QP through the bitrate control module); finally, performing entropy coding (such as CABAC or Huffman coding) on the quantized coefficients and other auxiliary information (such as motion vectors, prediction modes, etc.) to generate the final compressed bitstream. Throughout the encoding process, the internal bitrate control mechanism of the encoder will refer to the determined "video target bitrate", continuously adjust parameters such as the quantization level, and strive to make the size of the output bitstream approach the target value while ensuring the best possible preservation of video quality at the target resolution and frame rate.
[0089] Specifically, in step S700, the current conference video frame after adaptive optimization is transmitted. It should be understood that after a series of complex network state perception, video content analysis, target bitrate determination, and adaptive adjustment of encoding parameters (resolution, frame rate) based on this, the "current conference video frame after adaptive optimization" generated represents the most optimized visual data packet under the current constraints. Therefore, the current conference video frame after adaptive optimization is further transmitted. Transmission is the bridge connecting the intelligent decision-making at the sending end and the user experience at the receiving end. Its direct purpose is to efficiently and reliably (within the allowed real-time range) deliver the carefully encoded video data to one or more remote participants, enabling them to view the video images of the participants in real time, thereby achieving effective remote communication and collaboration.
[0090] In summary, the network adaptive video conference transmission optimization method according to the embodiments of the present application is elucidated. It systematically analyzes the historical data sequence of the network state to extract the deep trends and patterns of network behavior. At the same time, fully considering the different requirements of the current conference video frame type for transmission parameters (especially the bitrate), this key information is incorporated into the decision-making process. Through a specially designed fusion model, it can intelligently interact and collaborate on the historical evolution trend of the network state and the specific requirements of the current video content, so as to generate a video target bitrate that can not only adapt to future network fluctuations but also meet the current picture quality requirements. Based on this dynamically optimized target bitrate, the system further determines the appropriate target resolution and target frame rate, encodes and transmits the video. This method can effectively overcome the problems of poor adaptability and unreasonable resource allocation caused by traditional solutions that simply rely on the instantaneous network state or ignore video content differences. It aims to proactively ensure the smoothness and clarity of video conferences and improve the user experience in a complex and changing network environment.
[0091] Furthermore, a network adaptive video conference transmission optimization system is also provided.
[0092] Figure 5 For the block diagram of the network adaptive video conference transmission optimization system according to the embodiments of the present application. As Figure 5As shown, the network adaptive video conference transmission optimization system 500 according to an embodiment of the present application includes: a network status data acquisition module 510 that always obtains the time set of network status historical data from the sending end; a current conference video frame type acquisition module 520 for acquiring the type of the current conference video frame; a network status data preprocessing module 530 for preprocessing the time set of the network status historical data to obtain a sequence of network status historical data; a video target bitrate generation module 540 for determining a video target bitrate based on the type of the current conference video frame and the sequence of the network status historical data; a target resolution and frame rate determination module 550 for determining a target resolution and a target frame rate based on the video target bitrate and the type of the current conference video frame; a conference video frame encoding optimization module 560 for encoding the current conference video frame based on the target resolution and the target frame rate to obtain an adaptively optimized current conference video frame; and a conference video frame transmission module 570 that always performs data transmission on the adaptively optimized current conference video frame.
[0093] As described above, the network adaptive video conference transmission optimization system 500 according to an embodiment of the present application can be implemented in various wireless terminals, such as a server with a network adaptive video conference transmission optimization algorithm. In a possible implementation manner, the network adaptive video conference transmission optimization system 500 according to an embodiment of the present application can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the network adaptive video conference transmission optimization system 500 can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the network adaptive video conference transmission optimization system 500 can also be one of the many hardware modules of the wireless terminal.
[0094] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technologies in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.
Claims
1. A method for optimizing video conference transmission with network adaptability, characterized in that, including: a time set for obtaining historical network status data from a sending end; obtaining the type of the current conference video frame; preprocessing the time set of the historical network status data to obtain a sequence of historical network status data; determining a video target bitrate based on the type of the current conference video frame and the sequence of the historical network status data; determining a target resolution and a target frame rate based on the video target bitrate and the type of the current conference video frame; encoding the current conference video frame based on the target resolution and the target frame rate to obtain an adaptively optimized current conference video frame; performing data transmission on the adaptively optimized current conference video frame.
2. The network-adaptive video conference transmission optimization method according to claim 1, characterized in that The historical network status data includes an RTT mean value, a packet loss rate, a received bitrate, and a sent bitrate.
3. The network-adaptive video conference transmission optimization method according to claim 2, wherein Preprocessing the time set of the historical network status data to obtain a sequence of historical network status data includes: performing normalization processing and serialization processing on the time set of the historical network status data to obtain the sequence of the historical network status data.
4. The network-adaptive video conferencing transmission optimization method according to claim 3, characterized in that, Determining a video target bitrate based on the type of the current conference video frame and the sequence of the historical network status data includes: arranging the sequence of the historical network status data into a historical network status data time series matrix according to a network status sample dimension and a time dimension; extracting historical network status time series correlation features from the historical network status data time series matrix to obtain a historical network status data time series correlation feature vector; performing one-hot encoding on the type of the current conference video frame to obtain a current conference video frame type one-hot encoding vector; passing the historical network status data time series correlation feature vector and the current conference video frame type one-hot encoding vector through a network status-video type progressive complementary interaction perception network to obtain a network status-video conference type progressive interaction feature vector; performing feature decoding on the network status-video conference type progressive interaction feature vector to obtain a video target bitrate.
5. The network adaptive video conference transmission optimization method according to claim 4, wherein Extracting historical network status time series correlation features from the historical network status data time series matrix to obtain a historical network status data time series correlation feature vector includes: passing the historical network status data time series matrix through a historical network status time series correlation feature extractor based on a dilated convolutional neural network model to obtain the historical network status data time series correlation feature vector.
6. The network adaptive video conference transmission optimization method according to claim 5, wherein, Passing the historical network status data time series correlation feature vector and the current conference video frame type one-hot encoding vector through a network status-video type progressive complementary interaction perception network to obtain a network status-video conference type progressive interaction feature vector includes: performing multi-level hidden feature extraction on the historical network status data time series correlation feature vector and the current conference video frame type one-hot encoding vector to obtain a middle-level hidden feature encoding vector of the historical network status data time series, a middle-level hidden feature encoding vector of the current conference video frame type, a deep-level hidden feature encoding vector of the historical network status data time series, and a deep-level hidden feature encoding vector of the current conference video frame type; Based on the middle-level hidden feature encoding vectors in the time series of the network state historical data, the middle-level hidden feature encoding vectors of the current conference video frame type, the deep-level hidden feature encoding vectors in the time series of the network state historical data, and the deep-level hidden feature encoding vectors of the current conference video frame type, perform multi-level feature joint perception on the time series correlation feature vectors of the network state historical data and the one-hot encoding vectors of the current conference video frame type to obtain the progressive interaction feature vectors of the network state-video conference type.
7. The network-adaptive video conferencing transmission optimization method according to claim 6, wherein Based on the middle-level hidden feature encoding vectors in the time series of the network state historical data, the middle-level hidden feature encoding vectors of the current conference video frame type, the deep-level hidden feature encoding vectors in the time series of the network state historical data, and the deep-level hidden feature encoding vectors of the current conference video frame type, performing multi-level feature joint perception on the time series correlation feature vectors of the network state historical data and the one-hot encoding vectors of the current conference video frame type to obtain the progressive interaction feature vectors of the network state-video conference type includes: Perform low-level feature fusion on the time series correlation feature vectors of the network state historical data and the one-hot encoding vectors of the current conference video frame type to obtain the low-level fusion feature encoding vectors of the network state-video conference type; Perform high-level feature fusion on the middle-level hidden feature encoding vectors in the time series of the network state historical data and the middle-level hidden feature encoding vectors of the current conference video frame type to obtain the middle-level fusion feature encoding vectors of the network state-video conference type and the deep-level fusion feature encoding vectors of the network state-video conference type; Perform progressive complementary perception fusion on the low-level fusion feature encoding vectors of the network state-video conference type, the middle-level fusion feature encoding vectors of the network state-video conference type, and the deep-level fusion feature encoding vectors of the network state-video conference type to obtain the progressive interaction feature vectors of the network state-video conference type.
8. The network adaptive video conference transmission optimization method according to claim 7, characterized in that, Performing high-level feature fusion on the middle-level hidden feature encoding vectors in the time series of the network state historical data and the middle-level hidden feature encoding vectors of the current conference video frame type to obtain the middle-level fusion feature encoding vectors of the network state-video conference type and the deep-level fusion feature encoding vectors of the network state-video conference type includes: Perform middle-level feature fusion on the middle-level hidden feature encoding vectors in the time series of the network state historical data and the middle-level hidden feature encoding vectors of the current conference video frame type to obtain the middle-level fusion feature encoding vectors of the network state-video conference type; Perform deep feature fusion on the deep-level hidden feature encoding vectors in the time series of the network state historical data and the deep-level hidden feature encoding vectors of the current conference video frame type to obtain the deep-level fusion feature encoding vectors of the network state-video conference type.
9. The network self-adaptive video conference transmission optimization method according to claim 8, characterized in that, Performing feature decoding on the progressive interaction feature vectors of the network state-video conference type to obtain the video target bitrate includes: passing the progressive interaction feature vectors of the network state-video conference type through a target bitrate recommender based on a decoder to obtain the recommended decoded value of the video target bitrate.
10. A video conferencing transmission optimization system with network adaptability, characterized in that, Includes: A network status data acquisition module, which is used to obtain a time set of network status historical data from a sending end; A current conference video frame type acquisition module, which is used to obtain the type of the current conference video frame; A network status data preprocessing module, which is used to preprocess the time set of the network status historical data to obtain a sequence of network status historical data; A video target bitrate generation module, which is always used to determine a video target bitrate based on the type of the current conference video frame and the sequence of the network status historical data; A target resolution and frame rate determination module, which is used to determine a target resolution and a target frame rate based on the video target bitrate and the type of the current conference video frame; A conference video frame encoding optimization module, which is used to encode the current conference video frame based on the target resolution and the target frame rate to obtain an adaptively optimized current conference video frame; A conference video frame transmission module, which is used to perform data transmission on the adaptively optimized current conference video frame.
Citation Information
Patent Citations
Video conference method, system and device and storage medium
CN120050384A
Adaptive reward-driven video transport stream control method and device
CN120075486A
Video coding configuration method, system and device, and storage medium
WO2023134524A1
Cited By
Intelligent detection system for range hood
CN120252042A
Visual call information processing method and system based on 5G
CN120980184A
5g-based visual call information processing method and system
CN120980184B