Cross-layer optimization method and device for video stream exchange delay guarantee, and server

By dynamically adjusting the time window and traffic prediction model in the server kernel layer, optimizing the video streaming transmission rate, solving the problems of burstiness and response hysteresis in video streaming transmission, realizing the certainty of video stream exchange delay and improving the video playback quality.

CN120166077AActive Publication Date: 2025-06-17BEIJING UNIV OF POSTS & TELECOMM
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510637936.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

In the existing video streaming methods, the video application traffic is highly bursty and does not match the response speed of the media access control layer resource, resulting in poor video streaming quality, extended switching time, low video code rate and frequent switching of resolution.

Method used

In the kernel layer of the server, the input sequence length is determined by dynamically adjusting the strategy based on the time window, a time sequence is generated, and the traffic prediction model is input, and the video traffic prediction result data is output. Traffic shaping is performed based on the predicted result data, and the cross-layer transmission rate is optimized to ensure the certainty of the video stream switching delay.

Benefits of technology

The conflict between traffic burstiness and resource response hysteresis is solved, the reliability and stability of video streaming transmission is improved, the certainty of video streaming exchange delay is ensured, thus improving video streaming transmission efficiency, improving video playback quality and improving network resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120166077A_ABST
    Figure CN120166077A_ABST
Patent Text Reader

Abstract

The invention provides a cross-layer optimization method and device for video stream exchange delay guarantee and a server, and relates to the technical field of image communication, and the method comprises the steps: determining the length of an input sequence based on a time window dynamic adjustment strategy, generating a first time sequence, and obtaining a second time sequence according to a traffic prediction model; the first time sequence comprises a feature vector corresponding to each historical time point; the second time sequence comprises video traffic prediction result data corresponding to each target time point after the current moment; and extracting a third time sequence from the second time sequence according to the predicted time period, and performing traffic shaping to obtain cross-layer transmission rate optimization data for guaranteeing the video stream exchange delay so as to transmit the video clips in the application layer to a video stream request end through the time sensitive network. According to the invention, the reliability and stability of video stream transmission can be improved, and the certainty of video stream exchange time delay can be ensured, so that the video stream transmission efficiency and the video playing quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image communication technologies, and in particular to a cross-layer optimization method, apparatus, and server for ensuring video stream switching delay. Background Art

[0002] Video stream transmission not only requires a large amount of bandwidth resources but also extremely low (millisecond-level or even lower) switching delay guarantee. Therefore, the Time-Sensitive Networking (TSN) technology can be used to ensure the efficient transmission of video streams. For example, video data in the application layer of a streaming media server is exchanged via the time-sensitive network to an application program in a client device.

[0003] However, the existing video application traffic has strong burstiness and does not match the resource response speed of the media access control layer (MAC Layer), resulting in poor video stream transmission quality, specifically manifested as long switching delay, low video bit rate (low bandwidth utilization), and frequent resolution switching.

[0004] Therefore, there is an urgent need to design a method that can improve the reliability and smoothness of video stream transmission and ensure the certainty of video stream switching delay to improve the video stream transmission efficiency. Summary of the Invention

[0005] In view of this, embodiments of the present application provide a cross-layer optimization method, apparatus, and server for ensuring video stream switching delay to eliminate or improve one or more defects existing in the prior art.

[0006] One aspect of the present application provides a cross-layer optimization method for ensuring video stream switching delay, which is executed in the kernel layer of a server. The method includes: Determine the current input sequence length based on a time window dynamic adjustment strategy; generate a first time series sequence according to the input sequence length, and input the first time series sequence into a traffic prediction model so that the traffic prediction model outputs a second time series sequence; wherein, the first time series sequence includes feature vectors corresponding to respective historical time points; the feature vector is used to represent the context information corresponding to the video segment transmission process of transmitting a video segment in the application layer of the server to a video stream requester via the time-sensitive network at a corresponding historical time point; the second time series sequence includes: video traffic prediction result data corresponding to respective target time points after the current moment; According to the predicted time period, extract the video traffic prediction result data corresponding to multiple target time points from the second time series respectively to form a third time series, and perform traffic shaping on the third time series to obtain cross-layer transmission rate optimization data for ensuring the video stream switching delay. Based on the cross-layer transmission rate optimization data, at each of the target time points in the third time series, transmit the video segments in the application layer to the video stream request end via the time-sensitive network respectively.

[0007] In some embodiments of the present application, determining the current input sequence length based on the time window dynamic adjustment strategy includes: Obtain the first bitrate of the video segment transmission process corresponding to the nearest historical time point to the current moment, and the second bitrate of the video segment transmission process corresponding to another historical time point adjacent to the historical time point; Dynamically adjust the size of the target time window according to the bitrate change amplitude between the first bitrate and the second bitrate; Determine the current input sequence length according to the size of the target time window.

[0008] In some embodiments of the present application, the dynamically adjusting the size of the target time window according to the bitrate change amplitude between the first bitrate and the second bitrate includes: Calculate the size of the target time window according to the bitrate change amplitude between the first bitrate and the second bitrate, the pre-stored size of the current time window, the preset adjustment step coefficient, the sensitivity coefficient, the upper limit and the lower limit of the time window.

[0009] In some embodiments of the present application, performing traffic shaping on the third time series to obtain cross-layer transmission rate optimization data for ensuring the video stream switching delay includes: Perform a fast Fourier transform on the third time series to convert the video traffic prediction result data corresponding to each of the target time points in the third time series from the time domain to the frequency domain, and obtain the frequency domain data corresponding to each of the target time points in the third time series; Adopt a low-pass filter to perform filtering processing on each of the frequency domain data to obtain the filtered frequency domain data corresponding to each of the target time points in the third time series; Perform an inverse fast Fourier transform on each of the filtered frequency domain data to convert each of the filtered frequency domain data from the frequency domain back to the time domain, and obtain the traffic-shaped data corresponding to each of the target time points in the third time series; Obtain cross - layer transmission rate optimization data for ensuring video stream switching delay according to the traffic - shaped data corresponding to each of the target time points in the third time series.

[0010] In some embodiments of the present application, the traffic prediction model includes: a time - series prediction model based on the Transformer architecture.

[0011] In some embodiments of the present application, the obtaining cross - layer transmission rate optimization data for ensuring video stream switching delay according to the traffic - shaped data corresponding to each of the target time points in the third time series includes: Obtain the integral difference between the traffic - shaped data corresponding to each of the target time points in the third time series and the video traffic prediction result data corresponding to each of the target time points in the third time series; And, obtain the rate ratio corresponding to each of the target time points in the third time series according to the traffic - shaped data corresponding to each of the target time points in the third time series; Determine the target bandwidth rate corresponding to each of the target time points in the third time series according to the traffic - shaped data corresponding to each of the target time points in the third time series, the integral difference, the rate ratio corresponding to each of the target time points in the third time series, and a preset sampling interval, so that the target bandwidth rates corresponding to each of the target time points in the third time series constitute cross - layer transmission rate optimization data for ensuring video stream switching delay.

[0012] The second aspect of the present application provides a cross - layer optimization device for ensuring video stream switching delay, which is set in the kernel layer of the server. The device includes: A traffic predictor, configured to determine the current input sequence length based on a time - window dynamic adjustment strategy; generate a first time series according to the input sequence length, and input the first time series into a traffic prediction model so that the traffic prediction model outputs a second time series; wherein, the first time series includes feature vectors corresponding to each historical time point; the feature vector is used to represent the context information corresponding to the video segment transmission process of transmitting a video segment in the application layer of the server to a video stream request end via a time - sensitive network at a corresponding historical time point; the second time series includes: video traffic prediction result data corresponding to each target time point after the current moment; A time-aware traffic shaper is configured to extract video traffic prediction result data corresponding to multiple target time points from the second time series according to a predicted time period to form a third time series, and perform traffic shaping on the third time series to obtain cross-layer transmission rate optimization data for ensuring the video stream switching delay, so as to, based on the cross-layer transmission rate optimization data, transmit video segments in the application layer to the video stream request end via the time-sensitive network at each of the target time points in the third time series.

[0013] The third aspect of the present application provides a server. A cross-layer optimization device for ensuring video stream switching delay is provided in the kernel layer of the server; various video segments are stored in the application layer of the server; The cross-layer optimization device for ensuring video stream switching delay is configured to execute the cross-layer optimization method for ensuring video stream switching delay; The cross-layer optimization device for ensuring video stream switching delay is communicatively connected to a time-sensitive switch, and the time-sensitive switch is communicatively connected to a video stream request end, so that the cross-layer optimization device for ensuring video stream switching delay transmits video segments in the application layer to the video stream request end via the time-sensitive network.

[0014] In some embodiments of the present application, the video stream request end includes: an application program in an in-vehicle client device.

[0015] The fourth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the cross-layer optimization method for ensuring video stream switching delay is implemented.

[0016] The fifth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the cross-layer optimization method for ensuring video stream switching delay is implemented.

[0017] The sixth aspect of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the cross-layer optimization method for ensuring video stream switching delay is implemented.

[0018] The cross-layer optimization method for ensuring video stream switching delay provided by this application is executed in the kernel layer of the server. Based on the dynamic adjustment strategy of the time window, the current input sequence length is determined. A first time series sequence is generated according to the input sequence length, and the first time series sequence is input into a traffic prediction model so that the traffic prediction model outputs a second time series sequence. Among them, the first time series sequence includes feature vectors corresponding to respective historical time points. The feature vector is used to represent the context information corresponding to the video segment transmission process of transmitting the video segment in the application layer of the server to the video stream request end via the time-sensitive network at a corresponding historical time point. The second time series sequence includes video traffic prediction result data corresponding to respective target time points after the current moment. According to the prediction time period, video traffic prediction result data corresponding to multiple target time points are extracted from the second time series sequence to form a third time series sequence, and traffic shaping is performed on the third time series sequence to obtain cross-layer transmission rate optimization data for ensuring video stream switching delay. Based on the cross-layer transmission rate optimization data, at each of the target time points in the third time series sequence, the video segment in the application layer is transmitted to the video stream request end via the time-sensitive network, which can solve the conflict between traffic burstiness and resource response lag, improve the reliability and smoothness of video stream transmission, ensure the certainty of video stream switching delay, and further improve the video stream transmission efficiency, improve the video playback quality and increase the network resource utilization rate to ensure that users obtain a smoother viewing experience.

[0019] Additional advantages, objects, and features of this application will be partially described below and will become partially apparent to those of ordinary skill in the art after studying the following. Or they can be learned through the practice of this application. The objects and other advantages of this application can be achieved and obtained through the structure specifically pointed out in the specification and the drawings.

[0020] Those skilled in the art will understand that the objects and advantages that can be achieved by this application are not limited to the above specifically described, and the above and other objects that this application can achieve will be more clearly understood according to the following detailed description. Description of the Drawings

[0021] The drawings described herein are used to provide a further understanding of this application, form a part of this application, and do not limit this application. The components in the drawings are not drawn to scale, but only to illustrate the principle of this application. To facilitate the illustration and description of some parts of this application, the corresponding parts in the drawings may be enlarged, that is, they may become larger relative to other components in the exemplary device actually manufactured according to this application. In the drawings: Figure 1It is the first flow schematic diagram of the cross-layer optimization method for video stream switching delay guarantee in an embodiment of the present application.

[0022] Figure 2 It is the second flow schematic diagram of the cross-layer optimization method for video stream switching delay guarantee in an embodiment of the present application.

[0023] Figure 3 It is the third flow schematic diagram of the cross-layer optimization method for video stream switching delay guarantee in an embodiment of the present application.

[0024] Figure 4 It is the structural schematic diagram of the cross-layer optimization device for video stream switching delay guarantee in an embodiment of the present application.

[0025] Figure 5 It is the architecture schematic diagram among a server, a time-sensitive switch, and a video stream requester in an embodiment of the present application.

[0026] Figure 6 It is the architecture schematic diagram among a server, a time-sensitive switch, and an in-vehicle client device in a video in an application example of the present application.

[0027] Figure 7 It is the algorithm flowchart of the traffic predictor based on PatchTST in an application example of the present application. Detailed implementation manners

[0028] To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below in combination with the implementation manners and the drawings. Herein, the illustrative implementation manners of the present application and their descriptions are used to explain the present application, but do not limit the present application.

[0029] Herein, it also needs to be noted that in order to avoid obscuring the present application due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present application are shown in the drawings, while other details less related to the present application are omitted.

[0030] It should be emphasized that the term "including / containing" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0031] Herein, it also needs to be noted that if not otherwise specified, the term "connection" in this article can not only refer to a direct connection, but also represent an indirect connection with an intermediate.

[0032] In the following, embodiments of the present application will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0033] Time-Sensitive Networking (TSN) is a MAC layer technology defined by the IEEE 802.1 standard. It meets the transmission requirements of different flows through mechanisms such as high-precision clock synchronization (802.1AS), resource reservation (802.1 Qat), and traffic shaping (802.1Qav / Qbv). TSN technology can effectively coordinate the transmission of various data flows, ensuring that critical task data (such as autonomous driving data) and entertainment data can be efficiently and reliably transmitted on the same network, thereby improving the overall system performance and user experience.

[0034] Although time-sensitive networking technology can theoretically provide performance guarantees for autonomous driving, entertainment video applications, etc., in actual deployment, existing video stream transmission methods and time-sensitive network solutions often result in poor video playback quality. The main reason lies in the contradiction between the bursty traffic pattern at the application layer and the latency in MAC layer resource allocation. The specific problems are as follows: (1) The bursty transmission characteristics of video streams. Video streams have strong bursty transmission characteristics. The application layer divides video content into multiple video chunks and transmits all video chunks as instantaneously as possible. This bursty traffic pattern generates a large number of data packets in a short period of time without considering the bandwidth limitations and transmission constraints of the underlying network. As a result, it instantaneously consumes a large amount of bandwidth resources, increases transmission latency, and increases the risk of buffer overflow.

[0035] (2) Slow resource scheduling response. In a time-sensitive network, the MAC layer manages traffic through time slot scheduling and resource reservation mechanisms (such as Credit Based Scheduling, CBS). However, the MAC layer's response to bursty traffic is relatively slow. Since bursty traffic usually arrives simultaneously in a short period of time, the MAC layer may not be able to allocate sufficient resources for these data packets in a timely manner, resulting in the following problems: data packets are blocked while waiting for resources; when resources become available, the transmission resumes, and this latency may cause traffic transmission to lag; further leading to video playback stuttering and affecting the user experience.

[0036] Limitations of traditional methods: To solve the above problems, traditional methods usually reserve redundant resources for bursty application flows. For example, the 802.1 Qat protocol allocates resources by calculating the worst-case demand. However, this method has low resource utilization efficiency in low-traffic situations, thereby causing resource waste.

[0037] Based on this, in order to design a method that can improve the reliability and smoothness of video stream transmission and ensure the certainty of video stream exchange delay to improve the efficiency of video stream transmission, the embodiments of the present application respectively provide a cross-layer optimization method for video stream exchange delay guarantee, a cross-layer optimization device, system, physical device, computer-readable storage medium and computer program product for executing the cross-layer optimization method for video stream exchange delay guarantee, by adjusting the transmission timing of burst data packets so that they are scheduled when the MAC layer bandwidth resources are sufficient, thereby resolving the conflict between burst traffic and resource response.

[0038] The details are described in detail through the following examples.

[0039] Based on this, the embodiment of the present application provides a cross-layer optimization method for ensuring the delay of video stream exchange, which can be implemented by a cross-layer optimization device for ensuring the delay of video stream exchange, and is executed in the kernel layer of the server, see Figure 1 The cross-layer optimization method for ensuring the delay of video stream switching specifically includes the following contents: Step 100: Determine the current input sequence length based on the time window dynamic adjustment strategy; generate a first time series sequence according to the input sequence length, and input the first time series sequence into the traffic prediction model so that the traffic prediction model outputs a second time series sequence; wherein the first time series sequence includes feature vectors corresponding to each historical time point; the feature vector is used to represent the context information corresponding to the video segment transmission process of transmitting the video segment in the application layer of the server to the video stream request end via the time-sensitive network at a corresponding historical time point; the second time series sequence includes: video traffic prediction result data corresponding to each target time point after the current moment.

[0040] It should be noted that there is a strong correlation between video streams and transmission parameters (such as the sending window and buffer size), so its burst pattern can be learned in a data-driven manner. Therefore, in step 100 of the present application, a traffic prediction model is used that can output a time series consisting of video traffic prediction result data at each target time point in the future according to the time series consisting of context information at each historical time point, which can learn the dynamic mode of traffic forwarding from large-scale data, which is conducive to improving the accuracy of traffic change prediction.

[0041] However, for scenarios with high latency requirements such as in-vehicle video traffic prediction scenarios, since the traffic characteristics of this scenario have large dynamic fluctuations, if a fixed-length historical window is used as input processing, this static window design may lead to a decrease in prediction accuracy when facing drastic changes in video traffic.

[0042] Based on this, step 100 of the present application first determines the current input sequence length based on the time window dynamic adjustment strategy; it can be understood that the time window dynamic adjustment strategy is a way to achieve dynamic adjustment of the time window. Specifically, it can adjust the input window size according to the change of video bitrate. The purpose is to expand the window to obtain richer historical information when the traffic changes violently, and to shrink the window to improve the calculation efficiency when the traffic is stable.

[0043] At the same time, the time window dynamic adjustment strategy can also ensure that the traffic prediction model fully captures key historical information, and then can cooperate with the traffic shaping process in step 200 below to improve the accuracy of video stream transmission control.

[0044] Among them, the input sequence length refers to the length of the first sequence for which traffic prediction is to be performed currently. The first time sequence includes the feature vectors corresponding to each historical time point; the historical time point refers to the time point before the current moment, and the target time point refers to the time point after the current moment; in the embodiment of the present application, the time intervals between each historical time point and each target time point are the same, and the specific value can be set according to actual application requirements, and the present application does not make any limitations in this regard.

[0045] It can be understood that the feature vector corresponding to a historical time point is used to represent the context information corresponding to the video segment transmission process of transmitting the video segment in the application layer of the server to the video stream request end via the time-sensitive network at this historical time point. Among them, the feature vector can be defined as: Among them, represents the historical time point corresponding to this feature vector, that is, a step size; represents at the historical time point the video quality corresponding to the video segment transmission process of transmitting the video segment in the application layer of the server to the video stream request end via the time-sensitive network; represents at the historical time point the video bitrate corresponding to the video segment transmission process of transmitting the video segment in the application layer of the server to the video stream request end via the time-sensitive network; represents at the historical time point the buffer length corresponding to the video segment transmission process of transmitting the video segment in the application layer of the server to the video stream request end via the time-sensitive network; among them, is used to jointly reflect at the historical time point the context information corresponding to the video segment transmission process of transmitting the video segment in the application layer of the server to the video stream request end via the time-sensitive network; 。

[0046] The first time series sequence is the historical input sequence. If the current time window length contains N time durations, the first time series sequence X1 can be expressed as: Wherein, 、 to ( is )are the feature vectors corresponding to each historical time point respectively.

[0047] The second time series sequence is the predicted output sequence. The number of target time points in the second time series sequence is the same as the number of historical time points in the first time series sequence. The second time series sequence X2 can be expressed as: Wherein, 、 to are the video traffic prediction result data corresponding to each target time point respectively.

[0048] Step 200: According to the prediction time period, extract the video traffic prediction result data corresponding to multiple target time points from the second time series sequence respectively to form a third time series sequence, and perform traffic shaping on the third time series sequence to obtain cross-layer transmission rate optimization data for ensuring the video stream exchange delay. Based on the cross-layer transmission rate optimization data, at each of the target time points in the third time series sequence, transmit the video segments in the application layer to the video stream request end via the time-sensitive network respectively.

[0049] The prediction time period can be denoted as T, which is a constant. It can be understood that the number of target time points required by the prediction time period must be less than or equal to the number of target time points in the second time series sequence.

[0050] In step 200, the third time series sequence can be denoted as , which contains the video traffic prediction result data corresponding to multiple target time points extracted from the second time series sequence. The third time series sequence is used to represent the time-domain signal distribution before filtering. After performing traffic shaping processing on the third time series sequence, the traffic shaping data corresponding to each of the target time points in the third time series sequence is obtained. The sequence formed by the traffic shaping data corresponding to each of the target time points at this time can be denoted as ; then, the cross-layer transmission rate optimization data for ensuring the video stream exchange delay can be calculated according to the .

[0051] Among them, the cross-layer transmission rate optimization data for ensuring the video stream exchange delay can include the rate ratios corresponding to the respective target time points in the third time sequence. Furthermore, at each of the target time points in the third time sequence, according to the rate ratios corresponding to the respective target time points, in the kernel layer of the server, the video segments in the application layer of the server are transmitted to the video stream request end via the time-sensitive network.

[0052] As can be seen from the above description, the cross-layer optimization method for ensuring video stream exchange delay provided by the embodiments of the present application can solve the conflict between traffic burstiness and resource response hysteresis, improve the reliability and smoothness of video stream transmission, ensure the certainty of video stream exchange delay, and further improve the video stream transmission efficiency, improve the video playback quality and network resource utilization rate, so as to ensure that users obtain a smoother viewing experience.

[0053] In order to further improve the application effectiveness and reliability of the time window dynamic adjustment strategy, in a cross-layer optimization method for ensuring video stream exchange delay provided by the embodiments of the present application, refer to Figure 2 , step 100 in the cross-layer optimization method for ensuring video stream exchange delay specifically includes the following contents: Step 110: Obtain the first bit rate of the video segment transmission process corresponding to the nearest historical time point to the current moment, and the second bit rate of the video segment transmission process corresponding to another historical time point adjacent to this historical time point.

[0054] Step 120: Dynamically adjust the size of the target time window according to the bit rate change amplitude between the first bit rate and the second bit rate.

[0055] Step 130: Determine the current input sequence length according to the size of the target time window.

[0056] Step 140: Generate a first time sequence according to the input sequence length, and input the first time sequence into a traffic prediction model, so that the traffic prediction model outputs a second time sequence; wherein, the first time sequence includes the feature vectors corresponding to the respective historical time points; the feature vector is used to represent the context information corresponding to the video segment transmission process of transmitting the video segments in the application layer of the server to the video stream request end via the time-sensitive network at a corresponding historical time point; the second time sequence includes: the video traffic prediction result data corresponding to the respective target time points after the current moment.

[0057] To further improve the effectiveness and reliability of dynamically adjusting the size of the target time window according to the amplitude of the bitrate change between the first bitrate and the second bitrate, in a cross-layer optimization method for video stream switching delay guarantee provided in an embodiment of the present application, refer to Figure 3 Step 120 in the cross-layer optimization method for video stream switching delay guarantee specifically includes the following content: Step 121: Calculate the size of the target time window according to the amplitude of the bitrate change between the first bitrate and the second bitrate, the pre-stored size of the current time window, the preset adjustment step coefficient, the sensitivity coefficient, the upper and lower limits of the time window.

[0058] It can be understood that the current time window refers to the time window used for obtaining the length of the previous input sequence before the length of the current input sequence. Since at the time of executing Step 120, this time window is still the latest time window, it is therefore called the current time window. And the target time window refers to the time window used for obtaining the length of the current input sequence. After calculating the size of the target time window through Step 120, this target time window is the latest time window.

[0059] The size of the target time window The calculation formula is as follows: Wherein, represents the size of the current time window, represents the first bitrate of the video segment transmission process corresponding to the nearest historical time point to the current moment and the bitrate change amplitude between the second bitrate corresponding to another historical time point adjacent to this historical time point in the video segment transmission process; As the adjustment step coefficient, it is used to control the amplitude of the time window change; As the sensitivity coefficient, it is used to determine the response degree of the time window to the bitrate change; the tanh function is used to control: when the bitrate change amplitude is large, increase the time window size, when the bitrate is basically unchanged, the time window size remains unchanged, and the window size is adjusted smoothly. Finally, and are respectively used to control the lower and upper limits of the time window, restricting the range of the time window to prevent the calculation cost from increasing due to an overly large time window, and at the same time, it can also prevent the information obtained by the traffic prediction model from being insufficient due to an overly small time window.

[0060] To further improve the application effectiveness and reliability of the time window dynamic adjustment strategy, in a cross-layer optimization method for video stream switching delay guarantee provided in an embodiment of the present application, refer to Figure 2, step 200 in the cross-layer optimization method for ensuring video stream switching delay specifically includes the following content: Step 210: According to the predicted time period, extract the video traffic prediction result data corresponding to multiple target time points from the second time series to form a third time series.

[0061] It can be understood that extracting multiple target time points from the second time series can be carried out in the order of time series from front to back; for example, if the predicted time period T corresponds to two target time points, and the second time series contains 5 target time points The corresponding video traffic prediction result data for each ; then the third time series contains the target time points and The corresponding video traffic prediction result data for each .

[0062] Subsequently, in order to smooth the burst traffic of the application layer, the embodiments of the present application innovatively apply the Fast Fourier Transform (FFT) method and related filtering operations in the subsequent steps 220 and 230 to shape the video stream of the application layer. This method has two aspects of contribution: (1) Compared with the heuristic shaping method (such as Pacing), it effectively retains the original change trend of the traffic and adapts to the future dynamic changes of the video traffic; (2) Compared with the traditional machine learning-based shaping algorithm, it can quickly provide the shaping result and respond to the scheduling requirements of the system in a timely manner. That is: the following steps 220 and 230 first apply FFT to convert the future traffic predicted by the traffic prediction model from the time domain to the frequency domain. The key to smoothing lies in using a low-pass filter - it can remove the high-frequency components while retaining the low-frequency components, and these low-frequency components can capture the overall pattern of the traffic. Finally, apply the Inverse Fast Fourier Transform (iFFT) to convert the filtered low-frequency information back to the available predicted traffic time domain data.

[0063] Step 220: Perform a fast Fourier transform on the third time series to convert the video traffic prediction result data corresponding to each of the target time points in the third time series from the time domain to the frequency domain, and obtain the frequency domain data corresponding to each of the target time points in the third time series.

[0064] Step 230: Use a low-pass filter to perform filtering processing on each of the frequency domain data to obtain the filtered frequency domain data corresponding to each of the target time points in the third time series.

[0065] Step 240: Perform inverse fast Fourier transform on each of the filtered frequency-domain data to convert each of the filtered frequency-domain data from the frequency domain back to the time domain, obtaining the traffic-shaped data corresponding to each of the target time points in the third time series.

[0066] Step 250: Obtain cross-layer transmission rate optimization data for ensuring the video stream switching delay according to the traffic-shaped data corresponding to each of the target time points in the third time series.

[0067] Step 260: Based on the cross-layer transmission rate optimization data, at each of the target time points in the third time series, transmit the video segments in the application layer to the video stream request end via the time-sensitive network respectively.

[0068] However, in the above filtering process, the fast Fourier transform process requires sampling, and the traffic has the characteristics of burst and dynamic change, making it difficult to accurately capture its dynamic change trend. Therefore, in a cross-layer optimization method for ensuring video stream switching delay provided in an embodiment of the present application, the traffic prediction model includes: a time series prediction model based on the Transformer architecture, that is, the PatchTST model.

[0069] The PatchTST model is used to capture the context correlation in the time dimension. Although the PatchTST model has shown advantages in the above traffic prediction process, when directly applied to the in-vehicle video traffic prediction scenario, due to the large dynamic volatility of the traffic characteristics in this scenario, and the PatchTST model uses a fixed-length historical window as the input for processing, this static window design may lead to a decrease in prediction accuracy when facing the violently changing video traffic. In addition, when the PatchTST model is combined with the time series-aware shaper, the fixed window length may not be able to fully capture the key historical information, affecting the accuracy of the control strategy. Therefore, in step 100 of the present application, to solve this problem, the present application improves the input window mechanism of the PatchTST model so that it can adjust the input window size according to the video bitrate change, aiming to expand the window to obtain richer historical information when the traffic changes violently, and shrink the window to improve the calculation efficiency when the traffic is stable.

[0070] A Transformer is essentially an Encoder-Decoder architecture, which is divided into an encoding component and a decoding component. The encoding component consists of multiple layers of encoders, and each encoder contains two sub-layers: a Self-Attention layer and a feed-forward neural network layer. The decoding component also consists of multiple layers of decoders, and each decoder also contains two sub-layers: Self-Attention and a feed-forward neural network layer.

[0071] Initially, the first time series X1 is divided into multiple blocks and mapped to a high-dimensional space through an embedding mechanism. These embeddings serve as the input to the Transformer. In the Transformer, the self-attention mechanism plays a key role by capturing the global dependencies between different time periods. This is achieved based on the multi-head attention mechanism, which can be expressed as: (1) Where, is a trainable parameter, represents the concatenation of multiple attention heads; represents the multi-head attention mechanism for the input x; represents different attention heads.

[0072] Each individual attention head is calculated as: (2) Here, the function represents the self-attention mechanism, which allows the model to capture the dynamic correlations between time series and further process the time features by combining multiple layers of feed-forward networks. represents the video quality corresponding to the video segment transmission process of transmitting the video segment in the application layer of the server to the video stream requester via the time-sensitive network; represents the video bitrate corresponding to the video segment transmission process of transmitting the video segment in the application layer of the server to the video stream requester via the time-sensitive network; The buffer length corresponding to the video segment transmission process of transmitting the video segment in the application layer of the server to the video stream requester via the time-sensitive network; represents the trainable weight matrix for processing video quality features in the i-th attention head; represents the trainable weight matrix for processing video bitrate features in the i-th attention head; represents the trainable weight matrix for processing buffer length features in the i-th attention head.

[0073] Furthermore, in the filtering process of step 200, transmitting data packets using the obtained rate control may result in packet loss. This is because the low-pass filter removes high-frequency components, which substantially reduces the resources available for data transmission, thereby decreasing the total amount of transmitted data.

[0074] Based on this, to solve the above problems and ensure the normal transmission of the video stream and maintain the overall traffic size, the present invention application introduces adjustment measures based on the Fourier transform to maintain the integrity of the traffic. In a cross-layer optimization method for ensuring video stream switching delay provided in the embodiments of the present application, refer to Figure 3 , step 250 in the cross-layer optimization method for ensuring video stream switching delay specifically includes the following content: Step 251: Obtain the integral difference between the traffic-shaped data corresponding to each of the target time points in the third time series and the video traffic prediction result data corresponding to each of the target time points in the third time series.

[0075] Among them, the integral difference The calculation formula is as follows: Among them, is the third time series; is the sequence composed of the traffic-shaped data corresponding to each of the target time points; is a symbolic expression in differential and integral operations, used to describe the infinitesimal change of variable t or the integral dimension.

[0076] And, step 252: Obtain the rate ratio corresponding to each of the target time points in the third time series according to the traffic-shaped data corresponding to each of the target time points in the third time series.

[0077] The rate ratio The calculation formula is as follows: Step 253: Determine the target bandwidth rate corresponding to each of the target time points in the third time series according to the traffic-shaped data corresponding to each of the target time points in the third time series, the integral difference, the rate ratio corresponding to each of the target time points in the third time series, and a preset sampling interval, so that the target bandwidth rates corresponding to each of the target time points in the third time series constitute cross-layer transmission rate optimization data for ensuring video stream switching delay.

[0078] Among them, the compensated time-domain rate distribution formed by the target bandwidth rates corresponding to the respective target time points in the third time sequence has the following calculation formula: Among them, represents the sampling interval.

[0079] From a software perspective, the present application also provides a cross-layer optimization device for performing all or part of the video stream switching delay guarantee in the cross-layer optimization method for video stream switching delay guarantee, which is set in the kernel layer of the server. See Figure 4 The cross-layer optimization device for video stream switching delay guarantee specifically includes the following: A traffic predictor 10, configured to determine the current input sequence length based on a time window dynamic adjustment policy; generate a first time sequence according to the input sequence length, and input the first time sequence into a traffic prediction model, so that the traffic prediction model outputs a second time sequence; wherein, the first time sequence includes feature vectors corresponding to respective historical time points; the feature vectors are used to represent the context information corresponding to the video segment transmission process of transmitting a video segment in the application layer of the server to a video stream request end via a time-sensitive network at a corresponding historical time point; the second time sequence includes video traffic prediction result data corresponding to respective target time points after the current moment.

[0080] A time-aware traffic shaper 20, configured to extract video traffic prediction result data corresponding to multiple target time points from the second time sequence respectively according to a prediction time period to form a third time sequence, and perform traffic shaping on the third time sequence to obtain cross-layer transmission rate optimization data for guaranteeing video stream switching delay, so as to transmit the video segment in the application layer to the video stream request end via the time-sensitive network at each of the target time points in the third time sequence based on the cross-layer transmission rate optimization data.

[0081] The embodiments of the cross-layer optimization device for video stream switching delay guarantee provided by the present application can specifically be used to execute the processing flow of the embodiments of the cross-layer optimization method for video stream switching delay guarantee in the above embodiments, and its functions will not be elaborated here. Reference can be made to the detailed description of the embodiments of the cross-layer optimization method for video stream switching delay guarantee above.

[0082] The part of the cross-layer optimization for video stream switching delay guarantee by the cross-layer optimization device for video stream switching delay guarantee can be completed in the kernel layer of the server. The server can be a streaming media server, and the server is used to communicate with a client device via a time-sensitive network.

[0083] The above-mentioned client device may have a communication module (i.e., a communication unit), which can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.

[0084] Any suitable network protocol can be used for communication between the above-mentioned server and the client device, including network protocols that have not been developed as of the filing date of this application. The network protocol may, for example, include TCP / IP protocol, UDP / IP protocol, HTTP protocol, HTTPS protocol, etc. Of course, the network protocol may also include, for example, the RPC protocol (Remote Procedure Call Protocol) and the REST protocol (Representational State Transfer) used on top of the above-mentioned protocols.

[0085] As can be seen from the above description, the cross-layer optimization device for video stream switching delay guarantee provided in the embodiments of this application can solve the conflict between traffic burstiness and resource response latency, improve the reliability and smoothness of video stream transmission, guarantee the certainty of video stream switching delay, and further improve the video stream transmission efficiency, improve the video playback quality and increase the network resource utilization rate to ensure that users obtain a smoother viewing experience.

[0086] The embodiments of this application also provide a server. Refer to Figure 5 , this server can be a streaming media server, and a cross-layer optimization device for video stream switching delay guarantee is provided in the kernel layer of the server; various video segments are stored in the application layer of the server; The cross-layer optimization device for video stream switching delay guarantee is used to execute the cross-layer optimization method for video stream switching delay guarantee described in the embodiments of this application; The cross-layer optimization device for video stream switching delay guarantee is communicatively connected to a time-sensitive switch, and this time-sensitive switch is communicatively connected to a video stream request end, so that the cross-layer optimization device for video stream switching delay guarantee transmits the video segments in the application layer to the video stream request end via a time-sensitive network.

[0087] Among them, the video stream request end includes: an application program in an in-vehicle client device.

[0088] The server may include a processor, a memory, a receiver, and a transmitter. The processor and the memory may be connected through a bus or other means. Taking the connection through the bus as an example, the receiver may be connected to the processor and the memory in a wired or wireless manner. The processor may be a Central Processing Unit (CPU). The processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., or a combination of the above types of chips.

[0089] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as program instructions / modules corresponding to the cross-layer optimization method for video stream switching delay guarantee in the embodiments of the present application. By running the non-transitory software programs, instructions, and modules stored in the memory, the processor executes various functional applications and data processing of the processor, that is, implements the cross-layer optimization method for video stream switching delay guarantee in the above method embodiments.

[0090] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0091] The one or more modules are stored in the memory and, when executed by the processor, execute the cross-layer optimization method for video stream switching delay guarantee in the embodiments.

[0092] In some embodiments of the present application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, the memory, the receiver, and the transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to transmit and receive signals.

[0093] As an implementation manner, the functions of the receiver and the transmitter in the present application can be considered to be implemented by a transceiver circuit or a dedicated transceiver chip, and the processor can be considered to be implemented by a dedicated processing chip, a processing circuit or a general-purpose chip.

[0094] As another implementation manner, it can be considered to use a general-purpose computer to implement the server provided in the embodiments of the present application. That is, the program codes for implementing the functions of the processor, the receiver and the transmitter are stored in the memory, and the general-purpose processor implements the functions of the processor, the receiver and the transmitter by executing the codes in the memory.

[0095] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing cross-layer optimization method for ensuring the video stream switching delay are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium well-known in the technical field.

[0096] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the foregoing cross-layer optimization method for ensuring the video stream switching delay are implemented.

[0097] To further illustrate the above embodiments, the present application further provides a specific application example of the cross-layer optimization method for ensuring the video stream switching delay in an in-vehicle application scenario. First, the in-vehicle application scenario is described as follows: The popularization of entertainment systems in intelligent vehicles has greatly enriched the audio-visual experience of passengers, but at the same time has put forward higher requirements for the transmission performance of in-vehicle networks. With the accelerating transformation of the automotive industry towards intelligence, core technical fields represented by autonomous driving and advanced driver assistance systems (ADAS) are experiencing rapid development. In particular, the fully autonomous driving solution based on pure vision proposed by Tesla has attracted wide attention in the industry. At the same time, the growing demand of passengers for immersive entertainment experiences has prompted automobile manufacturers to widely deploy audio-visual devices in in-vehicle systems. For example, modern intelligent cockpits are usually equipped with multiple video screens for front and rear passengers and multi-channel audio players to significantly enhance the entertainment experience of passengers.

[0098] However, a large number of concurrent video streams impose a huge pressure on in-vehicle networks. These video streams not only consume a large amount of bandwidth resources but also pose extremely high requirements for transmission latency (usually within 10 milliseconds). In addition, in-vehicle networks also need to solve the problem of co-network transmission of audio and video streams and self-driving related streams (such as video analysis and radar point cloud data), and these self-driving streams have higher deterministic transmission requirements. Therefore, how to ensure the smooth transmission of multiple data streams within the same in-vehicle network and avoid mutual interference between streams has become a technical challenge. To solve this problem, in-vehicle networks have introduced Time-Sensitive Networking (TSN) technology.

[0099] See Figure 6 , the video stream requester can be an application in the client device, and this client device can be an in-vehicle client device. The in-vehicle client device first sends a video request to the server, prompting the server to select a target video segment. Then, the server returns a video segment with a suitable resolution according to the actual network link conditions. These video segments are transmitted through the time-sensitive network. In the above workflow, the method proposed in this application runs in the middleware (i.e., the kernel layer) of the streaming media server. Its goal is to coordinate the aggressive application layer video stream transmission and sluggish MAC layer resource allocation in the in-vehicle time-sensitive network. It smooths the transmission window before returning the video segment to the in-vehicle client device and adapts to the resource allocation of the underlying time-sensitive network. In Figure 6 , "GCL" represents the Gating List, which is used to configure the gating state of each port in the time-sensitive network during each time period, so as to control the forwarding window of a specific priority queue and ensure deterministic communication; "T0:oCooCoo" represents the gating states of different priority queues in time slice T0. Each letter represents the gating state of a queue (C represents Close, o represents Open). Therefore, this string sequence indicates that in time slice T0, the gates of priority queues 0, 2, 3, 5, 6 are open, allowing data to be sent, and the gates of priority queues 1, 4 are closed, not allowing data to be sent; "T1:CCooCoo" represents that in time slice T1, the gates of priority queues 2, 3, 5, 6 are open, allowing data to be sent, and the gates of priority queues 0, 1, 4 are closed, not allowing data to be sent.

[0100] In the above process, the method proposed by the application instance of this application needs to solve two problems.

[0101] (1) Accurately understand the application layer traffic pattern: There is a strong correlation between the video stream and transmission parameters (such as the send window and buffer size), so the burst pattern can be learned in a data-driven manner. The application example of this application designs a traffic predictor based on PatchTST. It learns the dynamic pattern of traffic forwarding from large-scale data based on the Transformer model, and uses the self-attention mechanism to effectively capture the long / short-term dependencies of historical traffic forwarding, which is beneficial to improving the accuracy of traffic change prediction.

[0102] (2) Driving the application layer traffic regulation to adapt to the MAC layer resource scheduling: Traditional methods use the Pacing method, but it will sacrifice the transmission efficiency of delay-sensitive data packets. Other learning-based traffic shaping algorithms cannot match the millisecond-level resource scheduling granularity of the MAC layer due to the inference delay (usually in seconds). In contrast, the application example of this application combines the Fast Fourier Transform (FFT) and related filtering operations to shape the application layer traffic. These methods are usually used in the field of digital signal processing and have not been applied to traffic shaping.

[0103] A cross-layer optimization method for video stream switching delay guarantee taking the in-vehicle application scenario as an example provided by the application example of this application analyzes the historical traffic behavior of the application layer, identifies its transmission mode, and adjusts the sending timing of burst data packets according to these mode characteristics to ensure scheduling when the MAC layer bandwidth resources are sufficient. The specific contents are as follows: (1) Traffic predictor based on the PatchTST model Feasibility of burst traffic prediction: From the perspective of network layer metrics (such as throughput and delay variation), the data packet transmission mode is highly dynamic and difficult to accurately predict. However, from the perspective of the sending behavior of the application layer, the traffic characteristics show a certain degree of identifiability and can be used for traffic prediction. During the DASH-based video stream transmission, application layer parameters such as the send window size, traffic trigger time, and buffer size are closely related, indicating that the burst pattern of the traffic mode can be learned using data-driven methods.

[0104] However, the traditional GRU model for traffic prediction mainly models short-term time series, and the prediction performance will significantly decline within the longer time window required by the time-series aware traffic shaper module. PatchTST, as a neural network model based on Transformer, has received extensive attention in recent years. It can efficiently learn the hidden dynamic pattern from large-scale datasets and construct an accurate time series prediction framework through the non-linear mapping of multi-layer neural networks. In addition, it uses the self-attention mechanism to effectively capture long-term and short-term dependencies, enabling it to better learn the traffic changes during video transmission.

[0105] See Figure 7 , the input for each step of the traffic predictor based on PatchTST is a feature vector , the traffic predictor based on PatchTST uses a sliding window mechanism for data preprocessing, generates a first time series, and captures the context correlation in the time dimension of the first time series through the PatchTST model to obtain a second time series.

[0106] Specifically, the layers connected in sequence of the PatchTST model are: an instance normalization layer, a Transformer encoder, and a smoothing and linear head layer. Among them, the instance normalization layer first normalizes the input data to ensure that the mean of each feature is zero and the standard deviation is one, thereby eliminating the influence between different scales and improving the stability of model training. Specifically, the input is a feature vector containing time series data, such as the resolution of a video stream , bit rate and buffer length , and these features will have different values at each time step. The instance normalization layer calculates the mean and standard deviation of the features at each time step and normalizes the data to zero mean and unit variance, and the output is the normalized feature vector. Next, the Transformer encoder processes these normalized data through the self-attention mechanism to capture the global and local dependencies between each time step in the time series. Specifically, the input of the Transformer is the normalized time series data (for example ), after being mapped to a high-dimensional space through the embedding layer, it is passed into the self-attention module. The self-attention mechanism calculates the relationship weights (query-key-value relationship) between each pair of elements in the input sequence and sums them up with weights to obtain a new representation for each time step, and the output is a feature matrix containing global and local information. The multi-head attention mechanism in the Transformer can calculate multiple relationships in parallel to capture information in different subspaces, while the positional encoding ensures that the time series information in the sequence is fully retained. Finally, the smoothing and linear head layer further processes the output of the Transformer encoder. First, it uses a low-pass filter to smooth the signal, filters out high-frequency fluctuations, and retains the long-term trend of the data, which can avoid short-term burst traffic interfering with bandwidth allocation. The smoothed data is then linearly transformed to output the final traffic prediction result, that is, the bandwidth requirement for each time step. These predicted results output are utilized by the time-aware traffic shaper to ensure the smooth and low-latency transmission of the video stream, thereby improving the quality of the video stream.

[0107] The core of the PatchTST model is based on the Transformer encoder implementation. The Transformer encoder consists of a multi-head attention mechanism layer, an instantiation layer, a feedback evaluation value layer, and an instantiation layer connected in sequence. Among them, for the multi-head attention mechanism layer, the input data comes from the feature matrix of the previous layer (the instantiation layer or the output of the previous Transformer encoder). These data are feature representations in the time series, such as the bit rate, resolution, buffer length, etc. of the video stream, and their shape is a matrix of (N, D), where N is the sequence length and D is the feature dimension of each time step. The input data is first mapped into query (Query), key (Key), and value (Value) matrices, and then the similarity between the query (Q) and the key (K) is calculated, and this similarity is used to weight and sum the value (V) to obtain the representation of each time step. The shape of the output is the same as the input, but the representation of each time step will contain more information, and this information comes from the weighted information of other time steps in the entire input sequence. By calculating multiple self-attention heads in parallel to capture information in different subspaces, the outputs of each head are finally concatenated together to form a complete representation. Next, the instantiation layer normalizes the output of the multi-head attention mechanism layer as its input. The normalization formula is: where represents the mean of the features, is the standard deviation of the features. Through this process, it is ensured that the features of each time step have zero mean and unit standard deviation in the dimension, which can eliminate the scale differences between different features and improve the consistency and training stability of the data. Then, the feedback evaluation value layer adopts a residual connection to combine the normalized data from the instantiation layer and the input data from the previous layer through the residual connection. The shape of the input is (N, D). The output data of the residual connection still has the shape of (N, D), and this data has considered the feedback information of the previous layer. This approach helps to alleviate the vanishing gradient problem in deep neural networks and promotes the smooth transmission of information. Finally, the second instantiation layer functions the same as the first instantiation layer, further normalizing the data to ensure that the distribution of each feature is stable and consistent, and further enhancing the stability and convergence during the training process. Through the processing of these four layers, the Transformer encoder can effectively capture the dynamic changes of video stream data, learn global and local dependencies, and generate high-quality traffic prediction results, thereby improving the prediction ability and stability of the entire PatchTST model.

[0108] Finally, the multi-head attention mechanism usually calculates the similarity between each query and key in parallel through multiple "heads", and then combines the results of these heads. This operation can help the model learn information from different subspaces, while Figure 6Where \(n_x\) represents the number of "heads" for such multiple parallel computations.

[0109] Although the PatchTST model has shown advantages in the above traffic prediction process, when directly applied to the in-vehicle video traffic prediction scenario, due to the large dynamic volatility of the traffic characteristics in this scenario, and PatchTST uses a fixed-length historical window as input for processing, this static window design may lead to a decrease in prediction accuracy when facing drastically changing video traffic. In addition, when PatchTST is combined with a temporal perception shaper, the fixed window length may not be able to fully capture key historical information, affecting the accuracy of the control strategy.

[0110] In the application example of this application, to solve this problem, the application example of this application improves the input window mechanism of PatchTST so that it can adjust the input window size according to the change of video bitrate. The purpose is to enable it to expand the window to obtain richer historical information when the traffic changes drastically, and to shrink the window to improve the calculation efficiency when the traffic is stable.

[0111] Through this technical means, the application example of this application overcomes the prediction error problem caused by the fixed window length, enabling PatchTST to more flexibly adapt to the dynamic changes of the video stream, thereby improving the prediction accuracy.

[0112] (2) Temporal Perception Traffic Shaper To smooth the burst traffic at the application layer, the application example of this application innovatively applies the FFT method and related filtering operations to shape the video stream at the application layer. The contributions of the temporal perception traffic shaper are reflected in two aspects: (1) Compared with heuristic shaping methods (such as Pacing), it effectively retains the original change trend of the traffic and adapts to the future dynamic changes of the video traffic; (2) Compared with traditional machine learning-based shaping algorithms, it can quickly provide shaping results and respond to the scheduling requirements of the system in a timely manner.

[0113] The temporal perception traffic shaper first applies FFT to transform the future traffic predicted by PatchTST from the time domain to the frequency domain. The key to smoothing lies in using a low-pass filter - it removes the high-frequency components while retaining the low-frequency components, and these low-frequency components can capture the overall pattern of the traffic. Finally, the Inverse Fast Fourier Transform (iFFT) is applied to convert the filtered low-frequency information back to the available predicted traffic time-domain data.

[0114] However, the following problems exist in the above filtering process: (1) FFT requires sampling, and the traffic has the characteristics of burst and dynamic changes, making it difficult to accurately capture its dynamic change trend; (2) Using the obtained rate control to transmit data packets may cause packet loss. This is because the low-pass filter removes the high-frequency components, substantially reducing the resources available for data transmission, thereby reducing the total amount of transmitted data. In the application example of this application, to solve the first problem, the traffic predictor based on PatchTST in Embodiment 1 is used to dynamically capture the data stream transmission mode, so as to provide traffic samples for the FFT method. To solve the second problem, that is, to ensure the normal transmission of the video stream and maintain the overall traffic size, the application example of this application introduces adjustment measures on the basis of the Fourier transform to maintain the integrity of the traffic. The specific operations are as follows.

[0115] To solve this mismatch, the time-domain waveforms before and after filtering are integrated, the difference between them is calculated, and interpolation is adjusted proportionally within the entire time domain. Then, the rate ratio of each target time point in the predicted time period T is calculated.

[0116] Through the above method, it can be ensured that the data stream is better shaped, and at the same time, it can be ensured that the overall traffic size can still be maintained after smoothing, and no packet loss will occur.

[0117] In summary, aiming at the contradiction between the bursty video traffic and the sluggish MAC layer resource allocation in the in-vehicle time-sensitive network, the application example of this application proposes a cross-layer optimization method for ensuring the deterministic video stream exchange delay to ensure the smooth transmission of the video stream and improve the playback quality. Specifically: 1) Traffic predictor based on PatchTST: Learn the dynamic change law of the video stream through the PatchTST traffic prediction model, accurately predict the future traffic demand, and provide accurate traffic data for the time-aware shaper.

[0118] 2) Time-aware shaper: Innovatively combine FFT and low-pass filter to smooth the video burst traffic at the application layer, and adjust the traffic transmission rate in real time according to the MAC layer resource scheduling requirements to ensure smooth transmission.

[0119] This method effectively solves the problem of mismatch between bursty video traffic and resource response in the in-vehicle time-sensitive network through the above mechanism, improves the video playback quality, and ensures that users obtain a smoother viewing experience.

[0120] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement it in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.

[0121] It should be clear that this application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of this application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of this application.

[0122] In this application, the features described and / or exemplified for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0123] The above are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, various changes and variations can be made to the embodiments of this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A cross-layer optimization method for ensuring delay in video stream switching, characterized in that: Executed in the kernel layer of the server, the method comprises: Determine the current input sequence length based on the time window dynamic adjustment strategy; generate a first time series sequence according to the input sequence length, input the first time series sequence into the traffic prediction model, so that the traffic prediction model outputs a second time series sequence; wherein, the first time series sequence includes feature vectors corresponding to each historical time point; the feature vector is used to represent the context information corresponding to the video segment transmission process of transmitting the video segment in the application layer of the server to the video stream request end via the time-sensitive network at a corresponding historical time point; the second time series sequence includes: video traffic prediction result data corresponding to each target time point after the current moment; According to the predicted time period, video traffic prediction result data corresponding to multiple target time points are extracted from the second timing sequence to form a third timing sequence, and traffic shaping is performed on the third timing sequence to obtain cross-layer transmission rate optimization data for ensuring the video stream exchange delay, so that at each target time point in the third timing sequence, the video clips in the application layer are respectively transmitted to the video stream request end via the time-sensitive network based on the cross-layer transmission rate optimization data.

2. The cross-layer optimization method for video stream exchange delay guarantee according to claim 1 is characterized in that: The determining of the current input sequence length based on the time window dynamic adjustment strategy includes: Obtaining a first bit rate of the video segment transmission process corresponding to a historical time point closest to the current moment, and a second bit rate of the video segment transmission process corresponding to another historical time point adjacent to the historical time point; Dynamically adjust the size of the target time window according to the bit rate change amplitude between the first bit rate and the second bit rate; The current input sequence length is determined according to the size of the target time window.

3. The cross-layer optimization method for video stream exchange delay guarantee according to claim 2 is characterized in that: The dynamically adjusting the size of the target time window according to the bit rate change amplitude between the first bit rate and the second bit rate includes: The size of the target time window is calculated according to the code rate change amplitude between the first code rate and the second code rate, the pre-stored size of the current time window, the preset adjustment step coefficient, the sensitivity coefficient, and the upper and lower limits of the time window.

4. The cross-layer optimization method for video stream exchange delay guarantee according to claim 1 is characterized in that: The performing traffic shaping on the third timing sequence to obtain cross-layer transmission rate optimization data for ensuring the video stream exchange delay includes: Performing a fast Fourier transform on the third time series to convert the video traffic prediction result data corresponding to each of the target time points in the third time series from the time domain to the frequency domain, and obtaining the frequency domain data corresponding to each of the target time points in the third time series; Using a low-pass filter to filter each of the frequency domain data to obtain filtered frequency domain data corresponding to each of the target time points in the third time series; Performing an inverse fast Fourier transform on each of the filtered frequency domain data to convert each of the filtered frequency domain data from the frequency domain back to the time domain, and obtaining the traffic shaped data corresponding to each of the target time points in the third time series; According to the traffic shaped data corresponding to each target time point in the third timing sequence, cross-layer transmission rate optimization data for ensuring the video stream switching delay is obtained.

5. The cross-layer optimization method for video stream exchange delay guarantee according to claim 4 is characterized in that: The traffic prediction model includes: a time series prediction model based on the Transformer architecture.

6. The cross-layer optimization method for video stream exchange delay guarantee according to claim 4 is characterized in that: The step of obtaining cross-layer transmission rate optimization data for ensuring video stream switching delay according to the traffic shaped data corresponding to each target time point in the third timing sequence includes: Obtaining an integral difference between the traffic shaped data corresponding to each of the target time points in the third time series and the video traffic prediction result data corresponding to each of the target time points in the third time series; And, according to the traffic shaped data corresponding to each of the target time points in the third time series, obtaining the rate ratio corresponding to each of the target time points in the third time series; According to the traffic shaped data corresponding to each target time point in the third timing sequence, the integrated difference, the rate ratio corresponding to each target time point in the third timing sequence, and the preset sampling interval, the target bandwidth rate corresponding to each target time point in the third timing sequence is determined, so that the target bandwidth rate corresponding to each target time point in the third timing sequence constitutes cross-layer transmission rate optimization data for ensuring video stream switching delay.

7. A cross-layer optimization device for ensuring delay in video stream switching, characterized in that: Set in the kernel layer of the server, the device includes: A traffic predictor is used to determine the current input sequence length based on a time window dynamic adjustment strategy; generate a first time series sequence according to the input sequence length, and input the first time series sequence into a traffic prediction model so that the traffic prediction model outputs a second time series sequence; wherein the first time series sequence includes feature vectors corresponding to each historical time point; the feature vector is used to represent context information corresponding to a video segment transmission process in which a video segment in the application layer of the server is transmitted to a video stream request end via a time-sensitive network at a corresponding historical time point; the second time series sequence includes: video traffic prediction result data corresponding to each target time point after the current moment; A timing-aware traffic shaper is used to extract video traffic prediction result data corresponding to multiple target time points from the second timing sequence according to a predicted time period to form a third timing sequence, and to perform traffic shaping on the third timing sequence to obtain cross-layer transmission rate optimization data for ensuring the video stream exchange delay, so as to transmit the video clips in the application layer to the video stream request end via the time-sensitive network at each target time point in the third timing sequence based on the cross-layer transmission rate optimization data.

8. A server, characterized in that: The core layer of the server is provided with a cross-layer optimization device for ensuring the delay of video stream exchange; the application layer of the server stores various video clips; The cross-layer optimization device for ensuring the delay of video stream switching is used to execute the cross-layer optimization method for ensuring the delay of video stream switching according to any one of claims 1 to 6; The cross-layer optimization device for video stream switching delay guarantee is communicatively connected to a time-sensitive switch, and the time-sensitive switch is communicatively connected to a video stream request end, so that the cross-layer optimization device for video stream switching delay guarantee transmits the video clips in the application layer to the video stream request end via a time-sensitive network.

9. The server according to claim 8, characterized in that: The video stream request end includes: an application program in a vehicle-mounted client device.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the cross-layer optimization method for ensuring the delay of video stream switching as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Video conference traffic prediction method and system based on time sequence representation learning

    CN113347384A

  • Variable bit rate service scheduling method based on flow prediction

    CN114374653A

  • Time-sensitive network data traffic prediction method and system, and storage medium

    CN114679388A

  • Mixed flow dynamic route planning method and device and computer equipment

    CN117135111A

  • Time delay optimization method and device for time-sensitive traffic

    CN118158091A