Hybrid transmission control method based on machine learning and heuristic algorithm

By combining machine learning and heuristic algorithms on mobile devices, the complexity and resource limitation problems in the prior art are solved, and efficient user experience quality optimization in dynamic network environments is achieved.

CN120050453APending Publication Date: 2025-05-27UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510103433.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing transmission methods are too complex to adapt to resource-constrained mobile device environments, and it is difficult to effectively optimize the quality of user experience (QoE) in dynamic network environments.

Method used

A hybrid transmission control method based on machine learning and heuristic algorithm is adopted, combined with the LSTM model and heuristic congestion control algorithm, by limiting the operating frequency of learning congestion control, reducing the calculation load, and dynamically adjusting the transmission priority according to the priority of the data block and the remaining time.

Benefits of technology

It achieves a balance of adaptability and performance in a dynamic network environment, reduces computing overhead, improves algorithm stability and effective utilization of computing resources, and significantly improves the quality of user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050453A_ABST
    Figure CN120050453A_ABST
Patent Text Reader

Abstract

The invention discloses a hybrid transmission control method based on machine learning and a heuristic algorithm, which comprises the following steps: S1, receiving data blocks generated by a mobile streaming media application program, discarding the data blocks missing the deadline, and selecting the data blocks according to the priority of the data blocks and storing the selected data blocks in a cache region; s2, judging whether the current moment is an integral multiple of the monitoring interval, if so, entering step S3, and otherwise, entering step S4; s3, acquiring network comprehensive data of the wireless network in the latest monitoring interval, inputting the network comprehensive data into the trained LSTM model, predicting to obtain a sending rate of a sent data block, and then entering the step S5; s4, according to the ACK returned by the application layer receiving end, a heuristic congestion control algorithm is adopted to calculate the sending rate of the sent data block, and then the step S5 is executed; and S5, reading the data blocks in the cache region, and sending the read data blocks to an application layer receiving end through a wireless network according to the sending rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to data transmission technology, and particularly to a hybrid transmission control method based on machine learning and heuristic algorithms. Background Art

[0002] With the rapid growth of self-media content and short video platforms, the demand for mobile streaming media transmission services by users is also increasing day by day. Users are not only content consumers but also producers. This transformation of roles makes mobile devices (such as smartphones, head-mounted displays, and portable cameras) key tools for users to generate and transmit high-quality content. Mobile streaming media services, especially emerging applications sensitive to latency such as augmented reality, virtual reality, and holographic communication, face challenges in a highly dynamic and unpredictable network environment, including bandwidth fluctuations, high latency, and large packet loss rates. These latency-sensitive services usually divide content into data chunks, and each data chunk contains multiple packets that must be transmitted on time. To ensure service continuity and smoothness, these data chunks are assigned specific priorities and deadlines to reflect the importance and urgency of the content they carry.

[0003] Traditional network transmission control schemes, such as heuristic congestion control (CC) algorithms (including Reno, Cubic, Vegas, and BBR, etc.), usually focus on global quality of service (QoS) metrics such as latency, throughput, and packet loss rate, aiming to maintain network stability and improve transmission efficiency. Further, congestion control algorithms based on machine learning (such as Remy, PCC, and Indigo, etc.) have emerged. They are designed to adapt to a highly dynamic network environment, especially wireless networks, to enhance QoS performance. However, these QoS metrics do not fully match the quality of experience (QoE) concerned by applications. For example, in the mobile streaming media scenario, even if a high throughput is achieved, if the video quality is poor or buffer interruptions occur frequently, it may not meet user needs. This is because the packets that make up each streaming media data chunk must be transmitted before a specific deadline, and for high-priority data chunks, failure to transmit on time will significantly reduce QoE.

[0004] To fill this gap, many studies have started to shift the focus from traditional QoS optimization to more complex QoE optimization, and proposed new methods including machine learning-based QoE-aware congestion control strategies (Floo), the combination of QoE-aware scheduling mechanisms and congestion control algorithms (DAP, D3T), etc. Floo trains a machine learning model that can switch the most suitable CC algorithm according to the current network state to optimize QoE. DAP introduces a chunk scheduler based on chunk priority and deadline to enhance QoE. However, the underlying CC algorithms of these solutions are still heuristic and difficult to adapt to the characteristics of dynamic mobile networks. D3T uses a machine learning-based CC algorithm in cooperation with a QoE-aware scheduler, providing an intelligent way to adapt to network dynamics, but also bringing challenges related to computational overhead. Especially for resource-limited mobile devices, running resource-intensive machine learning algorithms in real time may cause problems, leading to performance degradation and thus having a negative impact on the user experience. Therefore, it is necessary to minimize the algorithm complexity while maximizing the performance improvement effect of intelligent algorithms. Summary of the Invention

[0005] In view of the above deficiencies in the prior art, the hybrid transmission control method based on machine learning and heuristic algorithms provided by the present invention solves the problem that the existing transmission methods are too complex to adapt to the environment of resource-limited mobile devices.

[0006] To achieve the above invention objective, the technical solution adopted by the present invention is as follows:

[0007] Provide a hybrid transmission control method based on machine learning and heuristic algorithms, which includes the steps of:

[0008] S1. Receive the chunks generated by the mobile streaming media application, discard the chunks that miss the deadline, and select the chunks to be stored in the buffer according to the priority of the chunks;

[0009] S2. Determine whether the current moment is an integer multiple of the monitoring interval. If so, go to step S3; otherwise, go to step S4;

[0010] S3. Obtain the comprehensive network data of the wireless network within the most recent monitoring interval, input it into the trained LSTM model, predict the transmission rate of the chunks to be sent, and then go to step S5;

[0011] S4. Calculate the transmission rate of the chunks to be sent using the heuristic congestion control algorithm according to the ACK returned by the application layer receiver, and then go to step S5;

[0012] S5. Read the chunks in the buffer and send the read chunks to the application layer receiver through the wireless network according to the transmission rate.

[0013] Furthermore, the calculation method for the priority of data blocks includes:

[0014] S11. Calculate the remaining time required to complete the remaining packet transmissions of the data block based on the creation timestamp of the data block:

[0015]

[0016] Wherein, T rem and S rem are respectively the remaining time and the remaining size required to complete the remaining packet transmissions of the data block; T create is the creation timestamp of the data block; T current is the current timestamp; D is the deadline of the data block; R send is the sending rate; RTT last is the most recent round-trip delay;

[0017] S12. Determine whether the remaining time T rem is less than zero. If so, proceed to step S13; otherwise, proceed to step S14;

[0018] S13. Update the remaining time and then proceed to step S14, where e is the natural logarithm;

[0019] S14. Update the priority of the data block according to the remaining time T rem :

[0020] P new = T rem × (P max - P orig ) × (S rem / S)

[0021] Wherein, P new is the priority of the data block; P max is the set maximum priority value; P orig is the original priority of the data block; S is the volume of the data block.

[0022] Furthermore, the network comprehensive data is the average sending rate, average receiving rate, average round-trip delay, average estimated queuing delay, average estimated packet loss rate, and the number of packets in flight within the most recent monitoring interval.

[0023] Furthermore, the training method for the LSTM model includes:

[0024] S31. Obtain application trajectories and network trajectories under several multiple application scenarios as a trajectory dataset;

[0025] S32. Randomly select an unvisited trajectory from the trajectory dataset and run it in the simulator;

[0026] S33. Collect the network comprehensive data s of the simulator within the t-th monitoring interval t , and use the LSTM model and the expert strategy to obtain the predicted transmission rate value a t under the network comprehensive data s t and the optimal transmission rate

[0027] S34. Determine whether the selected trajectory has completed the entire trajectory. If so, go to step S36; otherwise, go to step S35;

[0028] S35. According to the current state of the LSTM model, select a t or input it into the simulator for execution, and store it in the total training dataset, update t = t + 1, and return to step S33;

[0029] S36. Use the total training dataset to train the LSTM model, and determine whether the LSTM model converges. If so, complete the training of the LSTM model; otherwise, return to step S32.

[0030] Furthermore, the LSTM model includes three states, namely 0, 1, and 2; when the state of the LSTM model is 0, select a t input it into the simulator for execution, and switch to state 1 with probability β in this state; when the state of the LSTM model is 1, select input it into the simulator for execution, and when the cumulative number of times of selecting in this state reaches 10 times, the state of the LSTM model switches to 2; when the state of the LSTM model is 2, select a t input it into the simulator for execution, and when the cumulative number of times of selecting a t in this state reaches 10 times, the state of the LSTM model switches to 0.

[0031] Furthermore, the probability β is updated once when the LSTM model is trained once using the total training dataset, and the update expression is:

[0032] β = β d

[0033] where d is the number of training times, and its initial value is 1.

[0034] Furthermore, the method for determining whether the LSTM model converges includes:

[0035] S361. Determine whether the probability β is less than or equal to a preset probability, and whether the difference in the quality of user experience after two consecutive LSTM model trainings is less than a preset value. If both are true, proceed to step S62; otherwise, the LSTM model has not converged.

[0036] S362. Determine whether the value of the loss function of the LSTM model is less than a preset loss. If so, the LSTM model has converged; otherwise, the LSTM model has not converged. The loss function is the mean square error function.

[0037] Furthermore, the expression for calculating the quality of user experience is:

[0038]

[0039] where QoE total is the total quality of user experience for N data blocks; N is the total number of data blocks included in the trajectory Φ(P n ) is the positive impact on the quality of user experience when data block n arrives on time; P n is the priority of data block n; m n is a constant. When all packets of data block n arrive on time, then m n = 1. When any packet in data block n fails to arrive on time, then m n = 0; ψ(P n ) is the negative impact on the quality of user experience when data block n fails to arrive on time.

[0040] Furthermore, the expert strategy is the Rate-M strategy, and its expression is α and f(t) are the scaling factor and the bandwidth at time t, respectively.

[0041] Furthermore, the monitoring interval is 30 - 50 ms.

[0042] Beneficial effects brought by the technical solution of the present invention:

[0043] 1. By combining heuristic congestion control and learning-based congestion control (i.e., the LSTM model), this solution effectively balances performance and complexity, enabling the mechanism to not only adapt to dynamic network environments but also operate efficiently on mobile devices with limited computing resources. By restricting the operating frequency of the learning-based congestion control (running the LSTM model once every monitoring interval), the computational load is reduced while maintaining the ability to quickly respond to network changes, achieving effective utilization of computing resources.

[0044] 2. Heuristic algorithms contribute to rapid decision-making, while learning-based algorithms optimize long-term decisions by analyzing historical data, enhancing adaptability across different time scales. Ultimately, this synergy improves algorithm stability and reduces the risk of overreaction or underreaction that may occur when relying solely on a single algorithm. Secondly, this two-tier design is computationally efficient and more practical than CC schemes that rely solely on machine learning, especially for deployment on mobile devices with limited computing resources. Generally, fully learning-based CC designs usually have potentially high computational overheads, and this scheme significantly reduces the computational overhead by invoking the learning-based CC algorithm at longer time intervals.

[0045] 3. When updating the priority of data blocks in this scheme, the remaining size of each data block is considered, so that data blocks with fewer remaining packets can be assigned higher priorities because they can complete transmission faster, thereby reducing the likelihood of missing deadlines. This scheme dynamically evaluates the transmission feasibility of each data block based on the remaining size, remaining time, and current network status. Combining this evaluation result with the original priority of the data block generates a recalculated transmission priority, and then the next data block to be transmitted is determined based on the updated priority, thus ensuring that the most urgent and important data blocks are sent first. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a flowchart of a hybrid transmission control method based on machine learning and heuristic algorithms.

[0047] Figure 2 It is a framework diagram of a system arranged for a hybrid transmission control method based on machine learning and heuristic algorithms.

[0048] Figure 3 It is a schematic diagram of the state transition mechanism of the LSTM model.

[0049] Figure 4 It is a schematic diagram of the performance under real-world tracking data; among them, (a) is a sports live broadcast application, a schematic diagram of the QoE performance under various network environments; (b) is a real-time game application, a schematic diagram of the QoE performance under various network environments; (c) is a movie video application, a schematic diagram of the QoE performance under various network environments.

[0050] Figure 5Schematic diagram for QoE performance evaluation at different monitoring intervals (MI), where (a) is the schematic diagram of the average QoE performance of each algorithm at different MIs; (b) is the cumulative distribution function (CDF) graph of the QoE performance of each algorithm in the full scenario when MI = 10 ms; (c) is the cumulative distribution function (CDF) graph of the QoE performance of each algorithm in the full scenario when MI = 30 ms; (d) is the cumulative distribution function (CDF) graph of the QoE performance of each algorithm in the full scenario when MI = 50 ms; (e) is the cumulative distribution function (CDF) graph of the QoE performance of each algorithm in the full scenario when MI = 70 ms; (f) is the cumulative distribution function (CDF) graph of the QoE performance of each algorithm in the full scenario when MI = 100 ms. Detailed implementation manners

[0051] The following describes the detailed implementation manners of the present invention to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

[0052] Refer to Figure 1 , Figure 1 shows a flowchart of a hybrid transmission control method based on machine learning and heuristic algorithms; as Figure 1 shown, this method S includes steps S1 to S5.

[0053] In step S1, receive data blocks generated by a mobile streaming media application program, discard data blocks that miss the deadline, and select data blocks according to the priority of the data blocks and store them in the buffer;

[0054] In an embodiment of the present invention, the calculation method of the priority of the data blocks includes:

[0055] S11. Calculate the remaining time required to complete the transmission of the remaining packets of the data block according to the creation timestamp of the data block:

[0056]

[0057] where, T rem and S rem are respectively the remaining time required to complete the transmission of the remaining packets of the data block and the remaining size; T create is the creation timestamp of the data block; T current is the current timestamp; D is the deadline of the data block; R send is the sending rate; RTT last is the most recent round-trip delay;

[0058] S12. Determine the remaining time T rem Whether it is less than zero. If so, go to step S13; otherwise, go to step S14:

[0059] S13. Update the remaining time Then go to step S14, where e is the natural logarithm;

[0060] S14. Update the priority of the data block according to the remaining time T rem :

[0061] P new = T rem ×(P max - P orig )×(S rem / S)

[0062] where P new is the priority of the data block; P max is the set maximum priority value; P orig is the original priority of the data block; S is the volume of the data block.

[0063] In step S2, determine whether the current moment is an integer multiple of the monitoring interval. If so, go to step S3; otherwise, go to step S4;

[0064] In step S3, obtain the network comprehensive data of the wireless network within the most recent monitoring interval, and input it into the trained LSTM model to predict the transmission rate of the data block to be sent. Then go to step S5.

[0065] In implementation, this solution preferably selects the network comprehensive data as the average transmission rate, average reception rate, average round-trip delay, average estimated queuing delay, average estimated packet loss rate, and the number of packets in flight within the most recent monitoring interval. Each time an ACK confirmation is received, Figure 2 the network system shown can directly calculate the network statistical information based on the information included in the ACK, that is, the reception rate, round-trip delay, estimated queuing delay, estimated packet loss rate, and the number of packets in flight. This part is a relatively mature technology in data transmission technology and will not be elaborated here.

[0066] In an embodiment of the present invention, the training method of the LSTM model includes:

[0067] S31. Obtain application trajectories and network trajectories in several multiple application scenarios as a trajectory dataset; the application trajectories can be sports live broadcasts, movies, games, and the network trajectories can be cellular networks in mobile scenarios such as walking, cycling, and cars.

[0068] S32. Randomly select an unvisited trajectory from the trajectory dataset and run it in the emulator; during the process of simulating the data transmission in the real world, the emulator simulates the behaviors of the sender and the receiver, sends data blocks according to the sending rate decided by the sender in real time, and updates the network state.

[0069] S33. Collect the network comprehensive data s of the emulator within the t-th monitoring interval t , and use the LSTM model and the expert strategy to obtain the predicted sending rate value a t of the network comprehensive data s t and the optimal sending rate The expert strategy is the Rate-M strategy. The Rate–M strategy directly adjusts the sending rate, and the formula is α and f(x) are the scaling factor and the bandwidth at time t respectively.

[0070] When selecting the expert strategy, this scheme selects three strategies: the CWND–M strategy, the CWND–A strategy, and the Rate–M strategy, and evaluates their QoE performance in all the collected trajectory data. The candidate strategy with the highest QoE performance will be selected as the expert strategy to guide the training of the learning model.

[0071] S34. Determine whether the selected trajectory has completed the entire trajectory. If so, go to step S36; otherwise, go to step S35.

[0072] S35. According to the current state of the LSTM model, select a t or and input it into the emulator for execution, store into the total training dataset, update t = t + 1, and return to step S33; every time the emulator receives the predicted rate value a t or the optimal sending rate , it takes it as the sending rate of the current time slot t and updates the network state at time t + 1.

[0073] As Figure 3 shown, in implementation, this scheme preferably sets the LSTM model to have 3 states, namely 0, 1, and 2; when the state of the LSTM model is 0, select a t and input it into the emulator for execution, and switch to state 1 with probability β in this state; when the state of the LSTM model is 1, select and input it into the emulator for execution. When the number of times of selecting in this state reaches 10 times, the state of the LSTM model switches to 2; when the state of the LSTM model is 2, select a t and input it into the emulator for execution. When the number of times of selecting a t in this state reaches 10 times, the state of the LSTM model switches to 0.

[0074] S36. Train the LSTM model using the total training dataset and determine whether the LSTM model converges. If it does, complete the training of the LSTM model; otherwise, return to step S32.

[0075] During the model training process, the optimal policy simultaneously minimizes the difference between the current LSTM model π and the expert policy π * Therein, l t is the loss function that measures the gap between the LSTM model and the expert policy. In this example, the mean squared error is used as the loss function, and the Adam optimizer is used to optimize the training process. The expression of the optimal policy is:

[0076]

[0077] Therein, is the expectation calculation of the random variable S under the state distribution d determined by the policy π π In this solution, the probability β is updated once when the LSTM model is trained once using the total training dataset. The update expression is: β = β

[0078] where d is the number of training times, and its initial value is 1. d

[0079] In implementation, the preferred method for determining whether the LSTM model converges in this solution includes:

[0080] S361. Determine whether the probability β is less than or equal to the preset probability and whether the difference in the quality of user experience after two adjacent LSTM model trainings is less than the preset value. If both are yes, go to step S62; otherwise, the LSTM model has not converged.

[0081] S362. Determine whether the value of the loss function of the LSTM model is less than the preset loss. If it is, the LSTM model has converged; otherwise, the LSTM model has not converged. The loss function is the mean squared error function.

[0082] The expression for calculating the quality of user experience is:

[0083]

[0084] Therein, QoE total is the total quality of user experience of N data blocks; N is the total number of data blocks included in the trajectory; Φ(P n ) is the positive impact on the quality of user experience when the data block n arrives on time; P n is the priority of the data block n; m n is a constant. When all packets of the data block n arrive on time, then mn = 1. When any packet in data block n fails to arrive on time, then m n = 0; ψ(P n ) is the negative impact on the quality of user experience when data block n fails to arrive on time.

[0085] In step S4, according to the ACK returned by the application layer receiver, the heuristic congestion control algorithm is used to calculate the transmission rate of the data block to be sent, and then step S5 is entered;

[0086] In step S5, the data block in the buffer is read, and according to the transmission rate, the read data block is sent to the application layer receiver through the wireless network.

[0087] The system diagram of the hybrid transmission control method based on machine learning and heuristic algorithm of this solution can be referred to Figure 2 , where the scheduler is used to execute step S1. The data blocks in step S1 are generated sequentially by the mobile streaming media application. Each data block has its priority, deadline and size, and these attributes are specified by the application. The Hybrid Congestion Control module based on heuristic and learning executes steps S3 and S4. After obtaining the transmission rate, it is transmitted in the form of packets under the management of the congestion control module.

[0088] In Figure 2 The right half of is the receiver of the application layer. At the receiver, if all the packets in the data block are received before the deadline, the block will pass through the filter and be delivered to the application layer. Otherwise, it will be discarded. After receiving the data block, the application layer sends the ACK to the sender through the feedback loop, notifying the system that the packet has been received and guiding further rate adjustment to optimize subsequent transmissions.

[0089] To facilitate the explanation of the effect of this solution, the performance of this solution will be verified by combining specific examples below:

[0090] Dataset: In a wireless network environment, a relatively comprehensive cellular network dataset is adopted. This dataset covers data collected under various traffic modes, such as cycling, walking, and taking buses, cars and trams. This dataset effectively reflects a wide range of speed ranges from slow walking to high-speed movement. The network state in the dataset is characterized by parameters such as bandwidth, propagation delay and random packet loss rate.

[0091] For the streaming tracking data of mobile applications, this example uses a set of video datasets, which includes video streaming chunks of three applications: movies, sports live broadcasts, and real-time games. This dataset records the generation timestamp, transmission deadline, as well as the size and priority information of each data chunk. The resolution of all videos is 1440P, and the frame rate is 30 frames per second.

[0092] Comparison algorithms: Considering that Qolit (the hybrid transmission control method of this solution) has scheduling (which is used to select data chunks according to their priorities and put them into the buffer) and congestion control (i.e., using the LSTM model and heuristic congestion control algorithm to calculate the sending rate), this embodiment selects the following combinations for transmission control to compare with Qolit.

[0093] CC algorithms: Use traditional heuristic algorithms and learning-based congestion control algorithms as benchmarks. For classic heuristic algorithms, this example compares a delay-based CC algorithm Vegas, two packet-loss-based CC algorithms Reno and Cubic, and a bandwidth-delay product-based CC algorithm BBR. For learning-based CC algorithms, this example selects a pre-trained pure LSTM model to ensure fair comparison.

[0094] Scheduling strategies: Considering the factors affecting QoE, this example adopts two direct scheduling strategies: earliest deadline first (EDF) and highest priority first (HPF). EDF preferentially transmits data chunks with more urgent deadlines, while HPF preferentially processes data chunks with higher priorities.

[0095] Figure 4 (a) shows the QoE performance of the sports live broadcast application in three mobile network environments, which are 1. walking and cycling, 2. bus and car, 3. train, corresponding to low-speed, medium-speed, and high-speed movements respectively; (b) shows the QoE performance of the real-time game application in three mobile network environments; (c) shows the QoE performance of the movie video application in three mobile network environments. In each figure, this solution combines Qolit with various heuristic algorithms (Reno, Cubic, BBR, Vegas) to obtain various variants of Qolit. For example, Qolit-Reno means using Reno as the heuristic algorithm in the hybrid control framework of Qolit and comparing it with pure heuristic algorithms.

[0096] Through Figure 4The results in [study area] show that various variants of Qolit have achieved significant improvements in QoE compared to heuristic algorithms. It can be observed that among the heuristic algorithms, the QoE performances of Reno and BBR are the worst and the best respectively. However, in the scenarios of three video content types (sports, movies, and games), the performance of Qolit-Reno is better than that of BBR, with average improvement rates of 12.7%, 11%, and 2.7% respectively. This improvement is attributed to the LSTM model in the Qolit hybrid CC control design, which learns how to correct inappropriate control behaviors during the training phase. This feature is particularly beneficial for Reno because Reno is prone to congestion misjudgment due to a high packet loss rate in the wireless environment, resulting in a reduced rate. LSTM can correct these errors in a timely manner, thus improving the performance of Reno. However, the QoE performance of Qolit-Reno is still lower than that of other Qolit variants.

[0097] When the control behaviors of other heuristic algorithms are better than Reno in terms of QoE performance, they require less correction by LSTM, so the overall QoE performance improves faster. This shows that combining Qolit with heuristic algorithms with better performance can obtain better QoE performance.

[0098] Figure 5 Schematic diagrams for QoE performance evaluation at different monitoring intervals (MI). This scheme selects two variants of Qolit combined with BBR and Vegas, and compares them with the pure LSTM algorithm. Among them, (a) shows the average QoE performance of each algorithm in the full scenario at different MIs; (b) shows the cumulative distribution function (CDF) graph of the QoE performance of each algorithm in the full scenario when MI = 10ms; (c) shows the cumulative distribution function (CDF) graph of the QoE performance of each algorithm in the full scenario when MI = 30ms; (d) shows the cumulative distribution function (CDF) graph of the QoE performance of each algorithm in the full scenario when MI = 50ms; (e) shows the cumulative distribution function (CDF) graph of the QoE performance of each algorithm in the full scenario when MI = 70ms; (f) shows the cumulative distribution function (CDF) graph of the QoE performance of each algorithm in the full scenario when MI = 100ms.

[0099] Through Figure 5The results show that different monitoring intervals (MI) have a significant impact on QoE performance. The Clean-LSTM model performs well at a smaller MI (10 ms), but its performance degrades as the MI increases. This indicates that a longer inference interval weakens the ability of Clean-LSTM to adapt to network state changes. In contrast, when the MI is set to 30 ms or higher, the performance of various Qolit variants is better than that of Clean-LSTM, and they reach their performance peaks at different MIs (e.g., Qolit-Vegas at 30 ms, Qolit-BBR at 50 ms). However, when the MI is lower than 30 ms, the performance of Qolit variants degrades because the control granularity of their two-layer mechanism is too similar, resulting in an interruption of control continuity, which is particularly evident in Qolit-BBR with a stronger closed-loop control logic.

[0100] Overall, the performance improvement of Qolit is attributed to the synergistic effect of learning-based CC and heuristic CC. Especially at higher MIs, the reduction in computational overhead enables this synergy to not only improve performance but also potentially outperform Clean-LSTM at the same or even lower MIs.

Claims

1. A hybrid transmission control method based on machine learning and heuristic algorithm, characterized in that: Includes steps: S1, receiving data blocks generated by a mobile streaming application, discarding data blocks that have missed deadlines, and selecting data blocks to store in a cache area according to the priority of the data blocks; S2, determine whether the current time is an integer multiple of the monitoring interval, if so, proceed to step S3, otherwise proceed to step S4; S3, obtaining the network comprehensive data of the wireless network in the most recent monitoring interval, and inputting it into the trained LSTM model to predict the sending rate of the sending data block, and then proceeding to step S5; S4, according to the ACK returned by the application layer receiving end, the heuristic congestion control algorithm is used to calculate the sending rate of the sending data block, and then enters step S5; S5. Read the data block in the buffer area, and send the read data block to the application layer receiving end through the wireless network according to the sending rate.

2. The hybrid transmission control method according to claim 1, characterized in that: The calculation method of the data block priority includes: S11. Calculate the remaining time required to complete the transmission of the remaining packets of the data block according to the creation timestamp of the data block: Among them, T rem and S rem T and T are the remaining time and size required to complete the transmission of the remaining packets of the data block respectively; create The creation timestamp of the data block; T current is the current timestamp; D is the expiration date of the data block; R send is the sending rate; RTT last is the most recent round trip delay; S12, determine the remaining time T rem Is it less than zero? If so, go to step S13; otherwise, go to step S14; S13. Update remaining time Then enter step S14, e is the natural logarithm; S14, according to the remaining time T rem , update the priority of the data block: P new =T rem ×(P max ―P orig )×(S rem / S) Among them, P new is the priority of the data block; P max is the maximum priority value set; P orig is the original priority of the data block; S is the size of the data block.

3. The hybrid transmission control method according to claim 1, characterized in that: The network comprehensive data includes the average sending rate, the average receiving rate, the average round trip delay, the average estimated queuing delay, the average estimated packet loss rate and the average number of in-flight packets in the most recent monitoring interval.

4. The hybrid transmission control method according to any one of claims 1 to 3, characterized in that: The training methods of LSTM model include: S31, obtaining application trajectories and network trajectories in multiple application scenarios as trajectory data sets; S32, randomly selecting an untraversed trajectory from the trajectory data set and inputting the trajectory into the simulator for execution; S33, collecting the network comprehensive data s of the simulator in the tth monitoring interval t , using LSTM model and expert strategy to obtain network comprehensive data s t The predicted value of the sending rate a under t and optimal sending rate S34, determine whether the selected trajectory has completed the entire trajectory, if so, proceed to step S36, otherwise proceed to step S35; S35. According to the current state of the LSTM model, select a t or Enter the simulator to execute. Store into the total training data set, update t=t+1, and return to step S33; S36. Use the total training data set to train the LSTM model and determine whether the LSTM model converges. If so, complete the training of the LSTM model, otherwise return to step S32.

5. The hybrid transmission control method according to claim 4, characterized in that: The LSTM model includes three states, namely 0, 1 and 2. When the state of the LSTM model is 0, select a t Input simulator execution, in this state switches to state 1 with probability β; when the state of the LSTM model is 1, select Enter the simulator to execute, select When the number of times reaches 10, the state of the LSTM model switches to 2; when the state of the LSTM model is 2, select a t Enter the simulator to execute, select a in this state t When the cumulative number of times reaches 10, the state of the LSTM model switches to 0.

6. The hybrid transmission control method according to claim 5, characterized in that: The probability β is updated once when the LSTM model is trained once using the total training data set, and the update expression is: β=β d Among them, d is the number of training times, and its initial value is 1.

7. The hybrid transmission control method according to claim 4, characterized in that: Methods for determining whether the LSTM model has converged include: S361, judging whether the probability β is less than or equal to the preset probability, and whether the difference in user experience quality after two consecutive LSTM model trainings is less than a preset value, if both are yes, proceeding to step S62, otherwise the LSTM model has not converged; S362. Determine whether the value of the loss function of the LSTM model is less than the preset loss. If so, the LSTM model has converged, otherwise the LSTM model has not converged; the loss function is a mean square error function.

8. The hybrid transmission control method according to claim 7, characterized in that: The expression for calculating user experience quality is: Among them, QoE total is the total user experience quality of N data blocks; N is the total data blocks included in the trace; Φ(P n ) is the positive impact of data block n arriving on time on the user experience quality; P n is the priority of data block n; m n is a constant. When all packets of data block n arrive on time, then m n = 1, when any packet in data block n fails to arrive on time, then m n =0;ψ(P n ) is the negative impact on user experience quality when data block n fails to arrive on time.

9. The hybrid transmission control method according to claim 7, characterized in that: The expert strategy is the Rate-M strategy, which is expressed as α and f(t) are the scaling factor and the bandwidth at time t respectively.

10. The hybrid transmission control method according to any one of claims 1-3 and 5-9, characterized in that: The monitoring interval is 30 to 50 ms.