A method and device for accelerating game networks based on reinforcement learning and multi-path collaboration
By employing a multi-path collaborative approach based on reinforcement learning, game data is segmented and error correction groups are identified. Combined with a network state prediction model, path weights are dynamically adjusted, solving the problems of high packet loss and high jitter in cross-border games and achieving stable and efficient data transmission.
Patent Information
- Application Number
- CN202510865569.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-06-26
AI Technical Summary
Existing technologies lack the ability to accurately predict and adaptively adjust real-time network conditions in cross-border online games, leading to problems such as high packet loss, high jitter, and wasted redundant bandwidth.
A multi-path collaborative approach based on reinforcement learning is adopted. By segmenting game data and adding error correction group identifiers, the path weights are dynamically adjusted using a network state prediction model to achieve multi-path allocation and forward error correction decoding, ensuring the stability and efficiency of data transmission.
In complex cross-border network environments, it dynamically reduces packet loss rate and latency jitter, and minimizes redundant bandwidth overhead to achieve stable and efficient game data transmission.
Smart Images

Figure CN120415649B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, and in particular to a method and apparatus for accelerating game networks based on reinforcement learning and multi-path collaboration. Background Technology
[0002] In cross-border online gaming scenarios, the connection between the client and overseas game servers typically involves complex public network routing. UDP packets are highly susceptible to factors such as link congestion, routing jitter, and ISP speed limits, leading to increased packet loss rates, fluctuating round-trip time (RTT), and aggravated transient jitter. To reduce perceived latency for players and improve data transmission reliability, various "game accelerator" products have emerged on the market. These products bypass some cross-border congestion sections by inserting proxy nodes between the user's terminal and the game server to build an encrypted tunnel.
[0003] With the increasing popularity of real-time competitive games and cloud gaming, users have placed higher demands on cross-border acceleration services: on the one hand, they need to maintain stable frame rates and instant response under high packet loss and high latency fluctuations; on the other hand, they want to reduce redundant overhead to save bandwidth when network conditions improve. Therefore, acceleration systems must have capabilities such as dynamic link quality assessment, on-demand redundant encoding and decoding, and multi-path concurrent scheduling.
[0004] In existing technologies, common solutions employ single-node or static multi-node proxy strategies, using fixed TCP acceleration channels or UDP forwarding channels coupled with simple retransmission mechanisms to reduce packet loss. Some solutions also introduce forward error correction (FEC) coding with fixed parameters or multi-substream transmission based on MPTCP, distributing the same data stream across several links to increase available bandwidth. However, these solutions typically operate with static redundancy ratios or static path weights, lacking fine-grained prediction and adaptive adjustment based on real-time network conditions.
[0005] The existing solutions mentioned above mainly have the following problems:
[0006] 1) Packet loss monitoring and FEC parameters cannot be dynamically adjusted according to real-time link fluctuations, resulting in excessive redundancy overhead under light load conditions and insufficient redundancy under heavy load conditions;
[0007] 2) Multi-path concurrency often distributes traffic using static round-robin or hash-based equal distribution methods, without combining real-time packet loss rate and node load for weight optimization, resulting in low link utilization.
[0008] 3) The lack of a forward-looking prediction mechanism for node quality leads to a lag in the switching between old and new nodes, which amplifies the network fluctuation window.
[0009] 4) The functional modules are often implemented separately, failing to form a closed loop of "monitoring-prediction-scheduling-error correction-feedback", making it difficult to continuously ensure transmission stability in complex cross-border network environments. Summary of the Invention
[0010] In view of this, embodiments of this application provide a game network acceleration method and apparatus based on reinforcement learning multi-path collaboration to solve the problem that the existing technology lacks an adaptive redundancy control and multi-path dynamic scheduling mechanism based on real-time network state prediction, which leads to high packet loss, high jitter and waste of redundant bandwidth in cross-border game data transmission.
[0011] A first aspect of this application provides a game network acceleration method based on reinforcement learning multi-path collaboration, comprising: on the client side, dividing the game data to be sent into fixed-length fragments, writing a continuous sequence number and an error correction group identifier for each fragment, and generating redundant fragments corresponding one-to-one with the fragments according to a preset forward error correction coding rule; sending probe data packets to at least two proxy nodes at a preset probe frequency, and determining the corresponding network monitoring vector based on the response data returned by each proxy node; inputting the network monitoring vector into a network state prediction model deployed on the proxy nodes or a control platform, and outputting a packet loss rate prediction result for a preset prediction window; and using the packet loss rate prediction result as a basis for further analysis. The packet rate prediction result serves as the state input for a reinforcement learning-based path scheduler. It calculates the path weights of at least two proxy nodes to obtain a multi-path allocation strategy. Following this strategy, fragments and their corresponding redundant fragments are sent in parallel between at least two proxy nodes, with the proxy node identifier used for transmission recorded in the packet header. At each proxy node, forward error correction decoding and reordering are performed on the received fragments and redundant fragments based on the error correction group identifier. This reconstructs the complete game data packet sequence and transmits it to the client side via a loopback link, enabling the client to receive the game data packet sequence and provide it to the local game process.
[0012] A second aspect of this application provides a game network acceleration device based on reinforcement learning multi-path collaboration, comprising: a fragmentation module, used to fragment game data to be sent on the client side into fixed-length fragments, write a continuous sequence number and an error correction group identifier to each fragment, and generate redundant fragments corresponding one-to-one with the fragments according to a preset forward error correction coding rule; a determination module, used to send probe data packets to at least two proxy nodes at a preset probe frequency, and determine the corresponding network monitoring vector based on the response data returned by each proxy node; and a prediction module, used to input the network monitoring vector into a network state prediction model deployed on the proxy nodes or a control platform, and output a packet loss rate prediction result for a preset prediction window; and calculate... The module is used to calculate the path weights of at least two proxy nodes and obtain a multi-path allocation strategy by taking the packet loss rate prediction result as the state input to the reinforcement learning-based path scheduler. The sending module is used to send fragments and corresponding redundant fragments in parallel between at least two proxy nodes according to the multi-path allocation strategy, and record the proxy node identifier used for sending in the packet header. The reconstruction module is used to perform forward error correction decoding and order rearrangement on the received fragments and redundant fragments at each proxy node according to the error correction group identifier, reconstruct the complete game data packet sequence, and transmit it to the client side through the loopback link so that the client side can receive the game data packet sequence and provide the game data packet sequence to the local game process.
[0013] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0014] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.
[0015] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects:
[0016] On the client side, the game data to be sent is fragmented into fixed-length segments. Each segment is written with a consecutive sequence number and an error correction group identifier, and redundant segments corresponding to each segment are generated according to a preset forward error correction coding rule. Probe data packets are sent to at least two proxy nodes at a preset probe frequency, and the corresponding network monitoring vector is determined based on the response data returned by each proxy node. The network monitoring vector is input into a network state prediction model deployed on the proxy nodes or control platform, and the packet loss rate prediction result for a preset prediction window is output. The packet loss rate prediction result is used as the state input to a reinforcement learning-based path scheduler to calculate the path weights of at least two proxy nodes and obtain a multi-path allocation strategy. According to the multi-path allocation strategy, the segments and the corresponding redundant segments are sent in parallel between at least two proxy nodes, and the proxy node identifier used for sending is recorded in the packet header. At each proxy node, forward error correction decoding and reordering are performed on the received segments and redundant segments according to the error correction group identifier to reconstruct the complete game data packet sequence, and the sequence is transmitted to the client side through a loopback link so that the client side receives the game data packet sequence and provides it to the local game process. This application can dynamically reduce packet loss rate and latency jitter in complex cross-border network environments, and minimize redundant bandwidth overhead while ensuring real-time performance, thereby achieving stable and efficient game data transmission. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the game network acceleration method based on reinforcement learning multi-path collaboration provided in an embodiment of this application;
[0019] Figure 2 This is a schematic diagram of the structure of the game network acceleration device based on reinforcement learning multi-path collaboration provided in the embodiments of this application;
[0020] Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0021] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0022] In existing technologies, most cross-border game accelerators adopt a single or static multi-node proxy strategy, and mitigate packet loss through fixed-parameter retransmission, forward error correction (FEC), or MPTCP traffic distribution. However, such solutions lack real-time prediction and adaptive adjustment of link quality: the redundancy ratio is difficult to reduce under light load, and redundancy compensation is insufficient under heavy load; multi-path traffic allocation usually relies only on static round-robin or hash mapping, and does not consider the combined impact of node load and instantaneous packet loss rate, resulting in high packet loss, high jitter, and redundant bandwidth waste problems that remain prominent.
[0023] To address the aforementioned shortcomings, the main technical challenge of this application is how to achieve real-time and stable transmission of game data in a complex and highly volatile cross-border network environment, while simultaneously considering packet loss suppression, latency control, and bandwidth efficiency.
[0024] To address this technical problem, this application provides a game network acceleration method based on reinforcement learning-based multi-path collaboration. The method first segments the game data and adds error correction group identifiers. A dynamic detection mechanism is used to obtain network monitoring vectors such as packet loss rate, round-trip latency, and jitter for each proxy node. Then, a network state prediction model deployed on the proxy nodes or control platform is used to prospectively estimate the packet loss rate within a preset prediction window. The prediction results are then used as state input to a deep reinforcement learning scheduler to calculate the path weights of each proxy node, forming a multi-path allocation strategy. Finally, the original segments and corresponding redundant segments are sent in parallel according to this strategy, and forward error correction decoding and reordering are performed at the proxy nodes before being sent back to the client.
[0025] The technical solution of this application mainly includes the following technical contents:
[0026] In the data encapsulation stage, this solution writes a globally consecutive number to each fragment and appends a unified group identifier, so that the encoding and subsequent recovery operations can accurately locate the fragment position. Then, based on configurable forward error correction parameters, one-to-one corresponding redundant fragments are generated in real time, realizing real-time protection of link quality at the sending end, rather than relying on the traditional approach of fixed redundancy ratio.
[0027] In terms of link quality assessment, the system uses a high-frequency detection mechanism to continuously collect packet loss rate, round-trip latency and jitter for each proxy node, combine them into a structured monitoring vector and update it on a rolling basis, thereby capturing network state changes within milliseconds, which is different from existing accelerators that only passively collect statistics at the session level.
[0028] The prediction module introduces a lightweight recursive network or sliding window model to infer the latest monitoring sequence and output the packet loss rate trend for the next short window. This forward-looking prediction makes up for the lag in traditional solutions that can only be adjusted after the fact.
[0029] At the scheduling level, a deep reinforcement learning network trained offline is used to map the prediction results, real-time latency, jitter, and node load together into action decisions, dynamically adjusting the traffic weights of multiple nodes to achieve fine-grained, closed-loop adaptive traffic distribution, rather than simple round-robin or static hash allocation.
[0030] During the data transmission phase, the system accurately splits the original and redundant fragments of the same error correction group according to the latest weight vector and writes the target node identifier in the packet header to ensure that the destination of each fragment is clear and controllable during multi-path transmission, significantly reducing unnecessary write-back and retransmission.
[0031] The receiving end aggregates and segments the data according to group identifiers. When the minimum decodeable threshold is reached, parallel decoding is triggered. The data is then rearranged according to the global sequence number and the entire frame is output. This "recovery upon arrival" strategy compresses the recovery latency to the millisecond level, which is in stark contrast to the traditional approach that relies on timeouts or batch aggregation before decoding.
[0032] Through the above technical solution, this application can proactively increase redundancy and reallocate traffic when network conditions deteriorate, and promptly reduce redundancy and converge to the optimal path when the network recovers, thereby achieving dynamic suppression of packet loss rate and latency jitter, and significantly reducing unnecessary bandwidth overhead, thus improving the stability and efficiency of cross-border game data transmission.
[0033] The technical solution of this application will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0034] Figure 1 This is a flowchart illustrating the game network acceleration method based on reinforcement learning multi-path collaboration provided in an embodiment of this application. Figure 1 As shown, this game network acceleration method based on reinforcement learning multi-path collaboration can specifically include:
[0035] S101, on the client side, the game data to be sent is divided into fixed-length fragments, each fragment is written with a continuous sequence number and an error correction group identifier, and redundant fragments corresponding to each fragment are generated according to the preset forward error correction coding rules.
[0036] S102, send probe data packets to at least two proxy nodes at a preset probe frequency, and determine the corresponding network monitoring vector based on the response data returned by each proxy node;
[0037] S103, input the network monitoring vector into the network status prediction model deployed on the agent node or control platform, and output the packet loss rate prediction result for the preset prediction window;
[0038] S104, using the packet loss rate prediction result as the state input to the reinforcement learning-based path scheduler, calculate the path weights of at least two agent nodes to obtain the multi-path allocation strategy.
[0039] S105, according to the multi-path allocation strategy, send the fragments and the redundant fragments corresponding to the fragments in parallel between at least two proxy nodes, and record the proxy node identifier used for sending in the packet header;
[0040] S106, at each agent node, forward error correction decoding and reordering are performed on the received fragments and redundant fragments according to the error correction group identifier, the complete game data packet sequence is reconstructed, and transmitted to the client side through the loopback link so that the client side can receive the game data packet sequence and provide the game data packet sequence to the local game process.
[0041] In some embodiments, a consecutive sequence number and an error correction group identifier are written to each fragment, and redundant fragments corresponding one-to-one with the fragments are generated according to a preset forward error correction coding rule, including:
[0042] According to the preset fragment length and error correction group capacity, the fragments to be sent are divided into several error correction groups, and a unified error correction group identifier is assigned to all fragments in the same error correction group.
[0043] Fragments within the same error correction group are assigned consecutive serial numbers in the order of their generation, and these consecutive serial numbers, along with the error correction group identifier, are written into the extended field of the fragment packet header.
[0044] Based on the preset forward error correction coding rules, linear block coding is performed on all segments within the same error correction group. Redundant segments corresponding to each segment are calculated in a finite field, and the corresponding consecutive sequence number and error correction group identifier are written into the packet header of each redundant segment.
[0045] Specifically, the client's send buffer invokes the fragmentation management module before data enters the network layer. The fragmentation management module performs a circular partitioning based on a preset fixed fragment length L_seg (e.g., 540 bytes, matching a typical MTU) and error correction group capacity parameters K_data and N_total (where K_data represents the number of original fragments in each group, and N_total represents the sum of original fragments and redundant fragments). The partitioning logic uses a circular queue, maintaining a group counter between the write and read pointers. A new error correction group ID (GID) is generated each time K_data original fragments are filled. All original fragments within the same group share this GID.
[0046] Furthermore, to ensure compatibility with existing UDP payloads, the system appends an 8-byte extension field before the application layer payload of each fragmented datagram, with the following structure:
[0047] 2-byte GID;
[0048] 4-byte consecutive sequence number SEQ (globally incremented from the start of the fragmentation management module, wrapping around after overflow);
[0049] The 2-byte fragment sequence number SID (0-N_total-1).
[0050] This extended field is placed at the very beginning of the application layer payload and does not affect the UDP header format.
[0051] Furthermore, after the fragments are generated, the fragment management module first writes a global consecutive sequence number SEQ to it, and then calculates the fragment sequence number SID in the group based on the number of original fragments already filled in the current error correction group; if the current number of fragments reaches K_data, the error correction coding logic is triggered.
[0052] Furthermore, the system's preset forward error correction coding adopts (N_total, K_data) linear block codes, with GF(256) selected for the finite field. The coding matrix G_enc is loaded from the configuration file at system startup, supporting both Reed-Solomon and RaptorQ matrix formats, and is mapped to a SIMD-friendly storage layout through a permutation index table. The encoding process performs row vector multiplication byte-by-byte on the K_data original fragments to obtain N_total-K_data redundant fragments; each redundant fragment maintains the same length and block identifier as its corresponding original fragment.
[0053] Furthermore, when a redundant fragment is generated, it is synchronously written with the same GID as its corresponding original fragment. The same auto-incrementing sequence is used in the SEQ field, and the SID field is filled with the value K_data up to N_total-1. This allows for direct differentiation of data categories at the receiving end while ensuring the integrity of the index within the group. After encoding, the original fragment and the redundant fragment enter the subsequent multipath scheduling process together.
[0054] For example, in a specific example, the following example uses the Reed-Solomon (12, 8) strategy with L_seg=540Byte, K_data=8, and N_total=12 to illustrate a complete sharding and redundancy generation process.
[0055] Assuming the game logic data block to be sent is 4320 bytes in size, the fragment management module sequentially extracts the data and cuts it into 8 original fragments, each fragment being 540 bytes long. The module assigns GID=1675 (decimal value for example) to these 8 fragments and assigns a global consecutive sequence number SEQ of 1024-1031 based on the generation order. The fragment sequence numbers SID within the group are written from 0 to 7 respectively.
[0056] After the 8th raw fragment is written, the fragment management module calls the FEC encoding engine. The encoding engine reads 8×540 bytes of data for the error correction group from the memory buffer and generates 4 redundant fragments at once through SIMD-parallel matrix operations.
[0057] The four generated redundant fragments are assigned the same GID=1675, and their global consecutive sequence numbers SEQ are 1032-1035 respectively. The fragment sequence numbers SID within the group are written as 8-11. The extended field of each redundant fragment occupies the first 8 bytes, consistent with the original fragment, followed by a 540-byte encoded data body.
[0058] At this point, a complete error correction group consists of 12 fragments, all of which enter the path scheduler's transmission queue and are assigned to different agent nodes based on subsequent reinforcement learning decisions. The receiving nodes use GIDs to classify the groups and SIDs to determine the decoding conditions. When the number of received fragments is greater than or equal to K_data, the inverse matrix can be solved and the missing fragments can be recovered. Then, the fragments are rearranged in SEQ order and output to the upper-layer reassembly buffer.
[0059] Through the above embodiments, the standardized encapsulation of fragments, sequence numbers, and error correction group identifiers is completed before the data enters the network for transmission. After the group is full, the corresponding redundant fragments are generated in real time. This allows the receiving end to recover the complete group with only K_data arbitrary fragments, which significantly improves the packet recovery success rate in cross-border links with high packet loss. At the same time, since the fragment length, group capacity, and encoding matrix can be flexibly configured as needed, the system can dynamically balance redundancy overhead and decoding latency while ensuring real-time performance, achieving stable, efficient, and bandwidth-optimized game data transmission.
[0060] In some embodiments, probe packets are sent to at least two proxy nodes at a preset probe frequency, and the corresponding network monitoring vector is determined based on the response data returned by each proxy node, including:
[0061] Construct a probe data packet and write the probe sequence number field, the sending timestamp field, and the integrity check field into the packet header;
[0062] According to the preset detection frequency, detection data packets are sent through detection channels corresponding to at least two proxy nodes respectively;
[0063] The client receives response data from each proxy node, counts the number of lost probe data packets based on the difference in probe sequence number, and calculates the round-trip delay based on the difference between the sending timestamp and the response arrival time.
[0064] Calculate round-trip time jitter using multiple consecutive response samples;
[0065] The packet loss rate, round-trip time, and round-trip time jitter obtained for the same proxy node are sequentially written into the network monitoring vector corresponding to the same proxy node.
[0066] Specifically, the client-side probe management module is responsible for constructing a fixed-length 96-byte probe data packet. The packet header contains a 2-byte probe sequence number field (SEQ), an 8-byte transmission timestamp field (TS), and a 4-byte integrity check field (HMAC). The remaining bytes are filled with preset constants. SEQ is incremented by an unsigned 16-bit counter, which restarts after wrapping around; TS uses a microsecond-level monotonically increasing timestamp; and HMAC is calculated using a session key pair (SEQ‖TS) to prevent replay attacks.
[0067] In some examples, the default probe frequency is 100 pkt / s. The probe management module establishes an independent UDP channel for each agent node, with the port number issued by the control platform. The module delivers a batch of probe packets to each channel every 10ms, and uses sendmmsg chaining at the system call level to reduce the number of context switches.
[0068] Each agent node immediately returns a new message containing only SEQ and TS upon receiving a probe data packet. The client side maintains two circular buffers, recording sent and received SEQ packets respectively. The module calculates the number of lost probe packets in the current period by comparing the difference n_lost = SEQ_send − SEQ_recv, and calculates the packet loss rate p_loss = n_lost / 100.
[0069] On the client side, when the response arrives, the current system time T_cur is taken, and the difference between this time and the corresponding TS is used to obtain the single-packet round-trip time RTT_i = T_cur − TS. A continuous set of 100 RTT_i constitutes a periodic sample set, and the average periodic round-trip time RTT_avg and the average jitter Jit = |RTT_i − RTT_{i−1}| are calculated.
[0070] The detection management module maintains a network monitoring vector V={p_loss,RTT_avg,Jit} for each agent node. This vector is refreshed upon completion of periodic statistics and passed to the network state prediction model via a shared memory queue. Vector updates employ a write-replacement strategy to ensure lock-free reading.
[0071] For example, in a specific case, assume the system simultaneously connects to proxy node A and proxy node B. The probe management module sends probe data packets with SEQ=0 to both nodes A and B at time 0s, and then cycles through these packets every 10ms. At 0.5s, node A loses 3 probe data packets due to link jitter; the module records n_lost_A=3 and calculates p_loss_A=3%. Node B experiences no packet loss within the same period, so p_loss_B=0%. For the 97 response data received by node A, RTT_avg_A=65ms and Jit_A=4ms are calculated; for node B, the corresponding RTT_avg_B=58ms and Jit_B=2ms. Finally, monitoring vectors V_A={0.03,65,4} and V_B={0,58,2} are generated and submitted to the prediction model.
[0072] By simultaneously conducting high-frequency probing of multiple proxy nodes at millisecond granularity on the client side and promptly calculating the network monitoring vector composed of packet loss rate, round-trip latency, and jitter, this embodiment achieves fine-grained quantitative evaluation of link quality, providing highly accurate input for subsequent network state prediction and reinforcement learning scheduling. It can trigger adaptive redundancy control and path weight adjustment in the early stages of network condition changes, thereby significantly reducing instantaneous packet loss and jitter in game data transmission.
[0073] In some embodiments, network monitoring vectors are input into a network state prediction model deployed on an agent node or control platform, and the output is a packet loss rate prediction result for a preset prediction window, including:
[0074] A training dataset is constructed based on historical network monitoring vectors and actual packet loss rates. The recurrent neural network model or sliding window regression model is then trained using the training dataset to obtain the parameters of the network state prediction model.
[0075] Load network state prediction model parameters into the agent node or control platform, and establish an inference interface for real-time invocation;
[0076] The latest network monitoring vectors from several consecutive periods are combined in chronological order to form an input tensor. Forward inference is performed through the inference interface to obtain the packet loss rate prediction value of each agent node within the preset prediction window.
[0077] The proxy node identifier and the corresponding packet loss rate prediction value are encapsulated into a packet loss rate prediction result, and the packet loss rate prediction result is provided to the reinforcement learning-based path scheduler.
[0078] Specifically, the system generates a network monitoring vector V every 1 second after completing a network probe on the client side. The vector includes three metrics: packet loss rate p_loss, average round-trip time RTT_avg, and latency jitter Jit. The monitoring link combines V with the actual packet loss rate y_true for the same period to form an entry.<ID,V,y_true> Write to the log. After the log is archived daily, the data cleaning module removes entries with missing fields and performs min-max normalization on the three metrics to obtain a continuous and balanced training dataset.
[0079] In some examples, the control platform can deploy two models, automatically selecting the appropriate one based on the node's computing power.
[0080] Recurrent Neural Network Model: Employs a two-layer LSTM with 64 hidden units per layer, an input window length of 10, and a prediction window length of 5 seconds. During training, an adaptive learning rate optimization algorithm is used, iterating until the validation set error converges.
[0081] Sliding window regression model: Enabled when node CPU is limited. The model calculates the feature mean of 5 continuous vectors, and then obtains the predicted value through ridge regression.
[0082] After the model training is completed, the system exports the weight file; the LSTM model exports the weight vector in ONNX format, and the regression model exports the weight vector in JSON format.
[0083] Furthermore, upon startup, the proxy node detects hardware capabilities, selects a matching model type, and downloads the latest weights. LSTM weights are quantized into an INT8 engine using TensorRT and loaded into the GPU or CPU SIMD instruction set; regression model weights are directly mapped to local memory. The inference service exposes a unified interface on a local socket, accepting only a fixed-length byte stream as input and returning a single floating-point number as the prediction result.
[0084] Furthermore, the inference management thread maintains a circular buffer B with a capacity of several vectors corresponding to the length of the model input window. When new monitoring vectors are written into B and fill the window, the thread concatenates the vectors in B in chronological order to form an input tensor, which is then submitted to the inference service via a local socket. The inference service returns the predicted packet loss rate y within the next 5 seconds. Subsequently, the thread encapsulates the node identifier PID and y into the <PID, y, t_pred> structure and writes it into the shared memory circular queue for the reinforcement learning scheduler to read.
[0085] For example, in a specific example, taking proxy nodes A and B as an example, the system sets the input window to 10 seconds and the prediction window to 5 seconds.
[0086] 1) At time t0, nodes A and B have each accumulated the last 10 monitoring vectors. The inference management thread concatenates these 10 vectors and submits them to the local inference service.
[0087] 2) Node A predicts a packet loss rate yA = 0.037, and node B predicts a packet loss rate yB = 0.012.
[0088] 3) The thread generates two records <PID = A, y = 0.037, t_pred = t0 + 5> and <PID = B, y = 0.012, t_pred = t0 + 5> and writes them into the shared queue.
[0089] 4) After reading, the reinforcement learning scheduler finds that the predicted value of node A is higher than the threshold of 0.03, and immediately increases the redundancy ratio and reduces the path weight; at the same time, it increases the weight of node B and diverts more data packets to node B to achieve preventive scheduling. The entire cycle takes about 40 ms, which is much shorter than the game frame interval and will not be perceived by the player.
[0090] By running the network state prediction model in real time on the proxy node or control platform, the system can obtain trend information in advance before the packet loss rate actually increases, leaving sufficient reaction time for the reinforcement learning scheduler. The scheduler can then adjust the redundancy and path weights in a timely manner, effectively reducing actual packet loss and jitter in a complex cross-border network environment, while avoiding redundancy waste when the network is in good condition, thus achieving a stable and efficient game data transmission experience.
[0091] In some embodiments, using the predicted packet loss rate as the state input to the reinforcement learning-based path scheduler, calculate the path weights of at least two proxy nodes respectively, and obtain a multi-path allocation policy, including:
[0092] Combine the predicted packet loss rate obtained for each proxy node with the corresponding real-time round-trip delay, delay jitter, and node processing load to form a state vector for reinforcement learning;
[0093] A deep reinforcement learning network, pre-trained offline and loaded with weight parameters, is used. The state vector is input into the deep reinforcement learning network, and the action value for each agent node is output.
[0094] An action is selected from a finite set of actions based on the action value. The action is used to adjust the path weights of each proxy node.
[0095] The path weights of at least two proxy nodes are updated based on the selected action to form a weight vector for all proxy nodes, and the weight vector is encapsulated into a multi-path allocation strategy.
[0096] Specifically, the path scheduler reads the latest predicted packet loss rate (ŷ), real-time round-trip time (RTT), latency jitter (Jit), and node CPU load (Load) from the shared circular queue every second, and concatenates these four metrics in a fixed order into a state vector S of length 4. If the number of currently active nodes is m, then this round of scheduling generates m state vectors. The scheduler then merges all vectors into a tensor S_all for network inference. To avoid scale differences affecting convergence, all four metrics are normalized online to the interval [0,1] according to the minimum-maximum coefficients saved during the initialization phase.
[0097] The scheduler uses a pre-trained deep Q-network. The network consists of three fully connected layers with a 4-dimensional input, 128 and 64 hidden layer nodes respectively, and an output dimension k equal to the size of the action set A. Each action in the action set A corresponds to a discrete instruction such as "increase or decrease the path weight of a node by 5%" or "remain unchanged". The network weights are trained using an ε-greedy strategy in a cloud simulation environment, then fixed into a fixed-point model file and distributed to all clients along with the version number.
[0098] The scheduler performs one network forward inference with S_all as input, obtaining k action values Q. Following a greedy principle, it selects the action a with the largest Q value; when multiple actions are equal, they are selected in ascending order of node ID. After executing action a, the path weight w_i of the corresponding node is immediately modified to ensure ∑w_i=1. To avoid drastic fluctuations, the system uses a smoothing coefficient α=0.2: new weight w_i′=α×w_i_new+(1−α)×w_i_old. The updated weight vector W={w_1′,…,w_m′} is encapsulated into a policy package and submitted to the sending queue management module via shared memory. The latter then allocates data fragments proportionally in the next scheduling cycle based on this package.
[0099] For example, in a specific case, assume three nodes, A, B, and C, are active simultaneously, with previous round weights of 0.4, 0.35, and 0.25, respectively. The current predicted packet loss rates (y) are 0.022, 0.041, and 0.010; RTTs are 60ms, 80ms, and 55ms; Jit times are 3ms, 6ms, and 2ms; and Load times are 0.30, 0.45, and 0.25. The scheduler constructs three state vectors and pushes them to the deep Q-network. The network outputs the Q-values corresponding to the action set, where the highest Q-value action 'a' is "reduce the weight of node B by 5% and increase the weight of node C by 5%". Executing 'a' generates temporary weights {0.4, 0.30, 0.30}; applying a smoothing coefficient α=0.2, the final weights {0.4, 0.33, 0.27} are obtained. The policy packet takes effect immediately after being written to the sending module, and data fragments are distributed according to the new ratio. The entire decision-making process takes approximately 20ms.
[0100] This embodiment incorporates the predicted packet loss rate, latency, jitter, and node load into the state vector, and uses a pre-trained deep Q-network to output actions in real time. This allows for proactive adjustment of traffic proportions before network quality deteriorates, enabling high-packet-loss nodes to quickly reduce their load and low-load nodes to take over in a timely manner. This significantly reduces the overall packet loss rate and latency jitter, while balancing bandwidth utilization and stability.
[0101] In some embodiments, according to a multi-path allocation strategy, fragments and corresponding redundant fragments are sent in parallel between at least two proxy nodes, and the proxy node identifiers used for transmission are recorded in the packet header, including:
[0102] Obtain the path weight vector for each proxy node based on the multi-path allocation strategy;
[0103] For each fragment and its corresponding redundant fragment belonging to the same error correction group, the distribution is carried out proportionally according to the path weight vector to determine the target proxy node for each fragment or redundant fragment.
[0104] Write the node identifier of the corresponding target proxy node into the header extension field of each fragment and redundant fragment to be sent, and encapsulate the node identifier together with the consecutive sequence number and the error correction group identifier.
[0105] The fragments and redundant fragments after writing the node identifier are added to the sending queue of the corresponding target proxy node and sent in parallel through independent transmission tunnels.
[0106] Specifically, the path scheduler has generated the updated weight vector W={w1, w2, ..., w...} in the previous embodiment. m}, where m is the number of currently active agent nodes, and ∑w i=1. Before entering a new round of data dequeueing, the sending module reads the latest W through shared memory and saves it to the local atomic variable table so that subsequent allocation logic can access it without locks.
[0107] After completing error correction coding, the data fragment manager generates a unified fragment object Pkt for each original fragment and its redundant fragments. The object contains: a global sequence number (SEQ), a group sequence number (SID), an error correction group identifier (GID), a length (LEN), a payload pointer (PTR), and a reserved target node field (PID). Pkt objects are managed in a circular buffer within a memory pool, and their dequeue order strictly follows the incrementing SEQ sequence to ensure predictable ordering of backend sending threads.
[0108] The sending module uses a cumulative weight interval method for proportional allocation. During the startup phase, a prefix sum is calculated on the weight vector W to obtain an interval array C, where C0 = 0, C1 = w1, C2 = w1 + w2, ..., C... m =1. Then, iterate through all Pkts within the same error correction group in the order they were generated, calculating a random number r∈[0,1) for each Pkt. Search for Pkt that satisfies C. j <r≤C j+1 The target node PID is determined as j+1 within the specified interval. This method naturally satisfies the proportional accuracy in the long-term statistical sense, while avoiding weight drift caused by integer division errors. If the business requires a more stringent real-time proportionality, random number generation can be replaced by a polling cumulative error method; however, this embodiment uses the interval method as an example.
[0109] After determining the PID, the sending module immediately writes three items to the extended field of the Pkt packet header: PID (2 bytes), SEQ (4 bytes), and GID (2 bytes). The write operation is completed directly through memory mapping, without copying the payload data. The written Pkt is immediately placed into the sending queue of the corresponding node. The sending queue is a lock-free doubly linked list, supporting multiple producers and consumers. Each agent node has an independent sending thread, which continuously retrieves Pkts in batches from its queue, using sendmmsg to send up to 32 fragments at a time, improving system call efficiency.
[0110] Each proxy node corresponds to an independent transport tunnel. If the operator supports multipath TCP, the tunnel uses MPTCP substreams; otherwise, DTLS encapsulates UDP. During the handshake phase, the tunnel completes key negotiation and writes the node's unique identifier (PID) to the session metadata. The sending thread maintains a long connection at the socket level and updates local statistics after each batch of data is sent for use by the probing module.
[0111] To ensure that different nodes do not block each other, the sending thread adopts a spin and back-off strategy. When the system buffer is full or congestion window shrinkage is detected, it will briefly self-sleep for 1 millisecond and then retry writing to avoid single-node congestion delaying global throughput. Before the thread exits, it will flush the socket to ensure that all remaining shards are sent out.
[0112] For example, in a specific example, assume that there are currently three active proxy nodes in the system: Node A, Node B, and Node C. The path weight vector W = {0.5, 0.3, 0.2}. The cumulative interval array C = {0, 0.5, 0.8, 1}. The current round of processing error correction group GID = 1800, which includes a total of twelve Packets, namely the original shards P0 to P7 and the redundant shards R8 to R11.
[0113] The sending module generates a random number 0.32 for P0 in sequence, which falls into the interval C0 < r ≤ C1, so PID = A;
[0114] The random number for P1 is 0.77, which falls into the C2 interval, and PID = B;
[0115] The random number for P2 is 0.11, and PID = A;
[0116] ……
[0117] Loop like this until R11. The statistical results show that the allocation ratio is approximately 5:4:3, which is close to the weight vector.
[0118] Subsequently, the Packets written with PID and other identifiers are classified into three sending queues according to PID. Queue A first takes 32 Packets and writes them out in batches; Threads B and C write out their respective batches in parallel. Because A has the highest weight, its thread frequency is the highest, but they do not affect each other. If the scheduler adjusts the weights in the next round, for example, A is reduced to 0.35 and C is increased to 0.35, only need to recalculate the C array and it will take effect immediately.
[0119] In this embodiment, through the cumulative interval algorithm combined with independent sending threads, high-precision proportional load distribution among multiple proxy nodes is achieved. Immediately writing the node identifier in the packet header ensures that the receiving end can accurately restore the flow direction and avoid relay path confusion. Independent transmission tunnels plus batch processing sending reduce the kernel context overhead and improve parallel throughput. At the same time, the congestion adaptive strategy of the sending thread prevents single-node blocking and ensures the smooth output of the overall data stream. Finally, this solution can still efficiently split traffic according to real-time weights in a dynamic network environment, and further cooperate with the aforementioned reinforcement learning scheduler to achieve end-to-end stable low-packet-loss game acceleration services.
[0120] In some embodiments, forward error correction decoding and sequential rearrangement are performed on the received shards and redundant shards according to the error correction group identifier to reconstruct a complete game data packet sequence, including:
[0121] At the proxy node, receive fragments and redundant fragments with the same error correction group identifier are written into the corresponding error correction group buffer, and an index is built according to the consecutive sequence number in the fragment.
[0122] When the number of fragments in the same error correction group buffer meets the preset decodeable threshold, the forward error correction decoding module is called to perform a finite field matrix inversion operation on the error correction group based on the preset forward error correction coding rules to restore the missing fragments.
[0123] After merging all the decoded fragments with the original received fragments, write them into the reordering buffer in consecutive sequence number order;
[0124] Determine whether consecutive sequence numbers in the reordering buffer are complete and consecutive. If they are complete, concatenate the corresponding fragment sets into a complete game data packet and output it to the feedback link.
[0125] Specifically, the proxy node maintains a hash table GTable in user space, mapping the error correction group buffer GBuffer to the error correction group identifier GID as the key. Internally, GBuffer uses a fixed-length slot array, with the number of slots equal to the encoding parameter N_total. Each slot stores a data pointer to the fragment and its arrival timestamp. Upon receiving a fragment, the packet header is parsed to obtain the GID, the group sequence number SID, and the consecutive sequence number SEQ. If the GBuffer corresponding to the GID does not exist, it is created and initialized immediately. Then, the fragment pointer is written to the slot according to the SID, and the SEQ is recorded in the slot's metadata for subsequent sorting by SEQ. This entire process uses lock-free write operations and relies on CAS to ensure concurrency safety.
[0126] The system configures a data segment K_data and a redundant segment N_total−K_data. When the number of valid slots in the GBuffer reaches K_data, the decoding threshold is met. To avoid prolonged resource consumption, the GBuffer also maintains a creation time T_start; if the current time minus T_start exceeds the maximum waiting time T_max, decoding is forcibly triggered even if the threshold has not been reached, ensuring real-time performance.
[0127] The node has a built-in FEC decoding library that supports two parameter levels: RS(12,8) and RS(40,20). When decoding is triggered, the module extracts the arrived fragments from the GBuffer in SID order and calls the kernel-mode high-performance SIMD instruction set to complete the finite field matrix inversion and missing fragment recovery. The recovered fragments are written to empty slots in the GBuffer and simultaneously written to the SEQ, at which point all twelve slots are filled.
[0128] The node maintains a circular reordering buffer RBuffer, writing data according to the global SEQ (sequence) incrementing sequentially. The write pointer WritePtr always points to the next position to be written. After decoding, the module copies the twelve fragments from GBuffer into RBuffer in ascending order of SEQ. If there are no gaps in consecutive SEQ in RBuffer, the splicing process is triggered; if a discontinuity is detected, the module pauses output and waits for the missing fragments to be filled in.
[0129] When the number of consecutive SEQs starting from LastOutputSeq+1 in RBuffer reaches a threshold M (e.g., 24 fragments), the module sequentially concatenates the contents of the corresponding fragments into a complete game data packet sequence PacketStream, and sends it back to the client via a loopback link. The loopback link uses DTLS encryption to ensure integrity. After output, LastOutputSeq is updated to the latest SEQ, and the RBuffer read pointer is synchronously moved forward to release space.
[0130] For example, in a specific instance, assuming RS(12,8) is currently used, the original fragments P0 to P7 and redundant fragments R8 to R11 within a certain error correction group GID=2457 arrive at the node via multiple parallel paths. The node first receives eight fragments: P0, P2, P3, P5, P6, R8, R9, and R10. After writing to the GBuffer, the effective slot number reaches K_data=8, and decoding immediately begins. The decoding module detects the missing four slots: P1, P4, P7, and R11. Based on the eight fragments that have arrived, it performs matrix inversion and successfully recovers the four missing fragments. After recovery, all twelve slots of the GBuffer are filled, and the module writes to the RBuffer in the order of SEQ45 to SEQ56. Since the previous error correction group has already been output and LastOutputSeq=44, there are no gaps in the RBuffer up to SEQ56. The system immediately concatenates 12×540 bytes of data to form a PacketStream and sends it back to the client. The entire process takes approximately 2ms, which is much less than the single-frame rendering interval.
[0131] By aggregating fragments by GID, decoding immediately upon meeting thresholds, and detecting the continuity of the reordering buffer, this embodiment achieves rapid recovery and sequential output of complete data even when packets arrive out of order and are partially lost. Combined with SIMD-optimized finite-field matrix inversion, decoding can be completed in milliseconds, ensuring real-time synchronization requirements for game scenes; simultaneously, a forced timeout decoding mechanism avoids long-tail latency caused by extreme packet loss. Ultimately, this process significantly reduces the impact of packet loss on the game experience in complex cross-border network environments, ensuring that the client receives a continuous, complete, and ordered sequence of game data packets.
[0132] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0133] Figure 2 This is a schematic diagram of the structure of the game network acceleration device based on reinforcement learning multi-path collaboration provided in an embodiment of this application. Figure 2 As shown, the game network acceleration device based on reinforcement learning multi-path collaboration includes:
[0134] The fragmentation module 201 is used to fragment the game data to be sent on the client side into fixed-length fragments, write a continuous sequence number and an error correction group identifier for each fragment, and generate redundant fragments that correspond one-to-one with the fragments according to the preset forward error correction coding rules.
[0135] The determination module 202 is used to send probe data packets to at least two proxy nodes at a preset probe frequency, and determine the corresponding network monitoring vector based on the response data returned by each proxy node.
[0136] Prediction module 203 is used to input network monitoring vectors into a network state prediction model deployed on a proxy node or control platform, and output packet loss rate prediction results for a preset prediction window.
[0137] The calculation module 204 is used to take the packet loss rate prediction result as the state input to the reinforcement learning-based path scheduler, calculate the path weights of at least two agent nodes, and obtain a multi-path allocation strategy.
[0138] The sending module 205 is used to send fragments and corresponding redundant fragments in parallel between at least two proxy nodes according to a multi-path allocation strategy, and to record the proxy node identifier used for sending in the packet header.
[0139] The reconstruction module 206 is used at each agent node to perform forward error correction decoding and reordering on the received fragments and redundant fragments according to the error correction group identifier, reconstruct the complete game data packet sequence, and transmit it to the client side through the loopback link so that the client side can receive the game data packet sequence and provide the game data packet sequence to the local game process.
[0140] In some embodiments, Figure 2The fragmentation module 201 divides the fragments to be sent into several error correction groups according to the preset fragment length and error correction group capacity, and assigns a unified error correction group identifier to all fragments in the same error correction group; it assigns consecutive sequence numbers to fragments in the same error correction group in the order of generation, and writes the consecutive sequence numbers and error correction group identifiers together into the extended field of the fragment packet header; based on the preset forward error correction coding rules, it performs linear block coding operation on all fragments in the same error correction group, calculates redundant fragments corresponding one-to-one with each fragment in a finite field, and writes the corresponding consecutive sequence number and error correction group identifier into the packet header of each redundant fragment.
[0141] In some embodiments, Figure 2 The determination module 202 constructs probe data packets and writes probe sequence number field, sending timestamp field, and integrity check field into the packet header; according to the preset probe frequency, probe data packets are sent through probe channels corresponding to at least two proxy nodes respectively; on the client side, response data from each proxy node is received, the number of lost probe data packets is counted according to the difference in probe sequence number, and the round-trip delay is calculated according to the difference between the sending timestamp and the response arrival time; the round-trip delay jitter value is calculated using multiple consecutive response samples; the packet loss rate, round-trip delay, and round-trip delay jitter obtained for the same proxy node are sequentially written into the network monitoring vector corresponding to the same proxy node.
[0142] In some embodiments, Figure 2 The prediction module 203 constructs a training dataset based on historical network monitoring vectors and actual packet loss rates, and uses the training dataset to train a recurrent neural network model or a sliding window regression model to obtain network state prediction model parameters. The network state prediction model parameters are loaded onto the agent nodes or control platform, and an inference interface is established for real-time calls. The latest continuous network monitoring vectors for several periods are combined in chronological order to form an input tensor, and forward inference is performed through the inference interface to obtain the packet loss rate prediction values for each agent node within a preset prediction window. The agent node identifier and the corresponding packet loss rate prediction value are encapsulated into a packet loss rate prediction result, and the packet loss rate prediction result is provided to the reinforcement learning-based path scheduler.
[0143] In some embodiments, Figure 2The computation module 204 combines the predicted packet loss rate for each agent node with the corresponding real-time round-trip latency, latency jitter, and node processing load to form a state vector for reinforcement learning. A deep reinforcement learning network, pre-trained offline and loaded with weight parameters, is used to input the state vector into the deep reinforcement learning network and output an action value for each agent node. An action is selected from a finite set of actions based on the action value, and the action is used to adjust the path weights of each agent node. The path weights of at least two agent nodes are updated based on the selected action to form a weight vector for all agent nodes, and the weight vector is encapsulated into a multi-path allocation strategy.
[0144] In some embodiments, Figure 2 The sending module 205 obtains the path weight vector for each proxy node according to the multi-path allocation strategy; for each fragment and corresponding redundant fragment belonging to the same error correction group, it allocates them proportionally according to the path weight vector to determine the target proxy node for each fragment or redundant fragment; it writes the node identifier of the corresponding target proxy node into the packet header extension field of each fragment and redundant fragment to be sent, and encapsulates the node identifier together with the consecutive sequence number and the error correction group identifier; it adds the fragment and redundant fragment after writing the node identifier to the sending queue of the corresponding target proxy node, and sends them in parallel through independent transmission tunnels.
[0145] In some embodiments, Figure 2 At the proxy node, the reconstruction module 206 writes received fragments and redundant fragments with the same error correction group identifier into the corresponding error correction group buffer and establishes an index according to the consecutive sequence number in the fragments. When the number of fragments in the same error correction group buffer meets the preset decodeable threshold, the forward error correction decoding module is called to perform a finite field matrix inversion operation on the error correction group based on the preset forward error correction coding rules to restore the missing fragments. After merging all the decoded fragments with the original received fragments, they are written into the reordering buffer in the order of consecutive sequence numbers. It is determined whether the consecutive sequence numbers in the reordering buffer are complete and consecutive. If they are complete, the corresponding fragment set is spliced into a complete game data packet and output to the feedback link.
[0146] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0147] Figure 3 This is a schematic diagram of the structure of the electronic device 3 provided in an embodiment of this application. Figure 3As shown, the electronic device 3 of this embodiment includes a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, it implements the steps in the various method embodiments described above. Alternatively, when the processor 301 executes the computer program 303, it implements the functions of each module / unit in the various device embodiments described above.
[0148] For example, computer program 303 may be divided into one or more modules / units, which are stored in memory 302 and executed by processor 301 to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of computer program 303 in electronic device 3.
[0149] Electronic device 3 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 3 may include, but is not limited to, processor 301 and memory 302. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device may also include input / output devices, network access devices, buses, etc.
[0150] Processor 301 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0151] The memory 302 can be an internal storage unit of the electronic device 3, such as a hard disk or RAM. The memory 302 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory 302 can include both internal and external storage units of the electronic device 3. The memory 302 is used to store computer programs and other programs and data required by the electronic device. The memory 302 can also be used to temporarily store data that has been output or will be output.
[0152] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0153] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0154] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0155] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. Multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of apparatus or units may be electrical, mechanical, or other forms.
[0156] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0157] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0158] If an integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium may include: any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0159] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although the technical solutions of this application are described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for accelerating game networks based on reinforcement learning and multi-path collaboration, characterized in that, include: On the client side, the game data to be sent is divided into fixed-length segments. A continuous sequence number and an error correction group identifier are written for each segment, and redundant segments corresponding to each segment are generated according to the preset forward error correction coding rules. Send probe data packets to at least two proxy nodes at a preset probe frequency, and determine the corresponding network monitoring vector based on the response data returned by each proxy node; The network monitoring vector is input into the network status prediction model deployed on the proxy node or control platform, and the packet loss rate prediction result is output for the preset prediction window. Using the packet loss rate prediction result as the state input to the reinforcement learning-based path scheduler, the path weights of the at least two proxy nodes are calculated to obtain a multi-path allocation strategy. According to the multi-path allocation strategy, the fragment and the redundant fragment corresponding to the fragment are sent in parallel between the at least two proxy nodes, and the proxy node identifier used for sending is recorded in the packet header; At each proxy node, forward error correction decoding and reordering are performed on the received fragments and redundant fragments according to the error correction group identifier to reconstruct the complete game data packet sequence. The sequence is then transmitted to the client side via a loopback link so that the client side can receive the game data packet sequence and provide it to the local game process.
2. The method according to claim 1, characterized in that, The step of writing a consecutive sequence number and an error correction group identifier to each fragment, and generating redundant fragments corresponding one-to-one with the fragments according to a preset forward error correction coding rule, includes: According to the preset fragment length and error correction group capacity, the fragment to be sent is divided into several error correction groups, and a unified error correction group identifier is assigned to all fragments in the same error correction group. Fragments within the same error correction group are assigned consecutive serial numbers in the order of their generation, and the consecutive serial numbers and the error correction group identifier are written together into the extended field of the fragment packet header; Based on the preset forward error correction coding rules, linear block coding is performed on all segments within the same error correction group. Redundant segments corresponding to each segment are calculated in a finite field, and the corresponding consecutive sequence number and error correction group identifier are written into the packet header of each redundant segment.
3. The method according to claim 1, characterized in that, The step of sending probe data packets to at least two proxy nodes at a preset probe frequency and determining the corresponding network monitoring vector based on the response data returned by each proxy node includes: Construct a probe data packet and write the probe sequence number field, the sending timestamp field, and the integrity check field into the packet header; The detection data packets are sent through detection channels corresponding to the at least two proxy nodes according to the preset detection frequency. The client receives response data from each proxy node, counts the number of lost probe data packets based on the difference in probe sequence number, and calculates the round-trip delay based on the difference between the sending timestamp and the response arrival time. Calculate round-trip time jitter using multiple consecutive response samples; The packet loss rate, round-trip time, and round-trip time jitter obtained for the same proxy node are sequentially written into the network monitoring vector corresponding to the same proxy node.
4. The method according to claim 1, characterized in that, The step of inputting the network monitoring vector into a network state prediction model deployed on a proxy node or control platform, and outputting a packet loss rate prediction result for a preset prediction window, includes: A training dataset is constructed based on historical network monitoring vectors and actual packet loss rates. The recurrent neural network model or sliding window regression model is then trained using the training dataset to obtain network state prediction model parameters. Load the network state prediction model parameters into the agent node or control platform, and establish an inference interface for real-time invocation; The latest network monitoring vectors from several consecutive periods are combined in chronological order to form an input tensor. Forward inference is performed through the inference interface to obtain the packet loss rate prediction value of each agent node within the preset prediction window. The proxy node identifier and the corresponding packet loss rate prediction value are encapsulated into a packet loss rate prediction result, and the packet loss rate prediction result is provided to the reinforcement learning-based path scheduler.
5. The method according to claim 4, characterized in that, The path scheduler, which uses the packet loss rate prediction result as state input and is based on reinforcement learning, calculates the path weights of each of the at least two proxy nodes to obtain a multi-path allocation strategy, including: The predicted packet loss rate obtained for each agent node is combined with the corresponding real-time round-trip latency, latency jitter and node processing load to form a state vector for reinforcement learning. A deep reinforcement learning network that has been pre-trained offline and loaded with weight parameters is used. The state vector is input into the deep reinforcement learning network, and the action value for each agent node is output. An action is selected from a finite set of actions based on the action value, and the action is used to adjust the path weight of each proxy node; The path weights of the at least two proxy nodes are updated based on the selected action to form a weight vector for all proxy nodes, and the weight vector is encapsulated into a multi-path allocation strategy.
6. The method according to claim 1, characterized in that, The step involves sending the fragment and its corresponding redundant fragment in parallel between at least two proxy nodes according to the multi-path allocation strategy, and recording the proxy node identifier used for transmission in the packet header, including: The path weight vector for each proxy node is obtained according to the multi-path allocation strategy; For each fragment and its corresponding redundant fragment belonging to the same error correction group, the target proxy node for each fragment or redundant fragment is determined by proportional allocation according to the path weight vector. Write the node identifier of the corresponding target proxy node into the header extension field of each fragment and redundant fragment to be sent, and encapsulate the node identifier together with the consecutive sequence number and the error correction group identifier. The fragments and redundant fragments after writing the node identifier are added to the sending queue of the corresponding target proxy node and sent in parallel through independent transmission tunnels.
7. The method according to claim 1, characterized in that, The step of performing forward error correction decoding and reordering on the received fragments and redundant fragments according to the error correction group identifier to reconstruct the complete game data packet sequence includes: At the proxy node, receive fragments and redundant fragments with the same error correction group identifier are written into the corresponding error correction group buffer, and an index is built according to the consecutive sequence number in the fragment. When the number of fragments in the same error correction group buffer meets the preset decodeable threshold, the forward error correction decoding module is called to perform a finite field matrix inversion operation on the error correction group based on the preset forward error correction coding rules to restore the missing fragments. After merging all the decoded fragments with the original received fragments, write them into the reordering buffer in consecutive sequence number order; Determine whether the consecutive sequence numbers in the reordering buffer are complete and consecutive. If they are complete, then concatenate the corresponding fragment sets into a complete game data packet and output it to the feedback link.
8. A game network acceleration device based on reinforcement learning multi-path collaboration, characterized in that, include: The sharding module is used to shard the game data to be sent on the client side into fixed-length shards, write a continuous sequence number and an error correction group identifier for each shard, and generate redundant shards that correspond one-to-one with the shards according to a preset forward error correction coding rule. The determination module is used to send probe data packets to at least two proxy nodes at a preset probe frequency, and determine the corresponding network monitoring vector based on the response data returned by each proxy node. The prediction module is used to input the network monitoring vector into the network state prediction model deployed on the agent node or control platform, and output the packet loss rate prediction result for the preset prediction window. The calculation module is used to take the packet loss rate prediction result as the state input to the reinforcement learning-based path scheduler, calculate the path weights of the at least two agent nodes, and obtain a multi-path allocation strategy. The sending module is configured to send the fragment and the redundant fragment corresponding to the fragment in parallel between the at least two proxy nodes according to the multi-path allocation strategy, and record the proxy node identifier used for sending in the packet header; The reconstruction module is used at each agent node to perform forward error correction decoding and reordering on the received fragments and redundant fragments according to the error correction group identifier, reconstruct the complete game data packet sequence, and transmit it to the client side through the loopback link so that the client side receives the game data packet sequence and provides the game data packet sequence to the local game process.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Heterogeneous encrypted multi-modal multimedia message real-time fragmentation transmission method for 5G network
CN120201420A
Data processing method and apparatuses, computer device, and storage medium
WO2023202243A1