An Adaptive Forward Error Correction Game Network Acceleration Method Based on Deep Reinforcement Learning

By employing a deep reinforcement learning-based adaptive forward error correction method, the problem of dynamically adjusting error correction parameters in cross-border network environments is solved, enabling efficient and stable transmission of game data packets, reducing packet loss rate and latency, and improving user experience.

CN120415648BActive Publication Date: 2025-11-14QINGFENG (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510865551.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-11-14
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Existing game accelerators cannot dynamically adjust forward error correction parameters and node switching in real time in cross-border network environments, resulting in unstable transmission quality and making it difficult to meet the needs of highly dynamic cross-border gaming scenarios.

Method used

An adaptive forward error correction method based on deep reinforcement learning is adopted. The network state vector is generated by periodic probing, and the error correction policy index is obtained by inputting it into the pre-trained model. The error correction parameter group is dynamically called for encoding and decoding, forming a closed loop of probing-decision-execution-feedback to achieve adaptive forward error correction.

Benefits of technology

It reduces packet loss rate, saves redundant bandwidth, maintains low-latency transmission, and improves node switching stability, significantly outperforming traditional accelerator solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120415648B_ABST
    Figure CN120415648B_ABST
Patent Text Reader

Abstract

This application provides an adaptive forward error correction method for game networks based on deep reinforcement learning. The method includes: periodically sending probe data packets to multiple candidate agent nodes and receiving echo data packets, generating a network state vector; inputting the network state vector into a pre-trained deep reinforcement learning model to obtain a forward error correction policy index; calling a target error correction parameter set and recording the current network environment data and the transmission performance of the target error correction parameter set back into a dynamic error correction policy library; performing forward error correction encoding on the game data packets to be transmitted according to the target error correction parameter set; decoding the received data packets and redundant packets at the agent node according to the target error correction parameter set, reconstructing missing game data packets, and forwarding the reconstructed game data packets to the target game server or game client for use in game network acceleration transmission. This application can reduce packet loss rate, save redundant bandwidth, maintain low-latency transmission, and enhance node switching stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technology, and in particular to an adaptive forward error correction game network acceleration method based on deep reinforcement learning. Background Technology

[0002] In cross-border scenarios, online games typically use the User Datagram Protocol (UDP) for real-time interaction. Due to the complexity of cross-border backbone link routing, intense bandwidth competition, and differences in export operator strategies, domestic players are prone to packet loss, latency fluctuations, and increased jitter when accessing overseas game servers. To mitigate these issues, the industry widely deploys "game accelerators," which improve transmission stability through dedicated tunnels, proximity access nodes, and data packet redundancy mechanisms.

[0003] Under conditions of random and uncontrollable fluctuations in cross-border link quality, game transmission needs to simultaneously achieve low latency and packet loss recoverability. An ideal solution should be able to: have fine-grained awareness of real-time network conditions; achieve an adaptive balance between packet loss, congestion, and bandwidth overhead; and quickly switch to a better transmission path when network conditions change abruptly.

[0004] Existing accelerators are mainly optimized in three aspects: first, by monitoring link quality through fixed probe intervals; second, by activating forward error correction at a preset ratio when packet loss exceeds a threshold; and third, by selecting proxy nodes based on simple judgment rules. However, such solutions generally treat probe, error correction, and node switching as independent functional modules, lacking a unified decision-making mechanism, and the error correction parameters are usually triggered by static thresholds, making it difficult to adapt to dynamic network changes.

[0005] However, the detection results only make immediate threshold judgments and do not form continuous state variables, which cannot support fine-grained error correction decisions. Fixed-proportion forward error correction consumes additional bandwidth when the link is good, but may not be enough to recover data when the link deteriorates, lacking adaptive capability. Node switching and error correction initiation are independent of each other, and packet loss during switching cannot be recovered in time by the previous error correction strategy, which can easily lead to additional data loss. Summary of the Invention

[0006] In view of this, embodiments of this application provide an adaptive forward error correction game network acceleration method based on deep reinforcement learning to solve the problem that the existing technology cannot dynamically adjust forward error correction parameters and node switching based on real-time network status, making it difficult to meet the transmission requirements of highly dynamic cross-border game scenarios.

[0007] A first aspect of this application provides an adaptive forward error correction game network acceleration method based on deep reinforcement learning, comprising: periodically sending probe data packets to multiple candidate agent nodes and receiving echo data packets; generating a network state vector based on the echo data packets; inputting the network state vector into a pre-trained deep reinforcement learning model to obtain a forward error correction policy index, wherein the forward error correction policy index indicates a target error correction parameter group in a dynamic error correction policy library or indicates disabling forward error correction; calling the target error correction parameter group from the dynamic error correction policy library according to the forward error correction policy index, and recording the current network environment data and the transmission performance of the target error correction parameter group back to the dynamic error correction policy library; performing forward error correction encoding on the game data packets to be transmitted according to the target error correction parameter group to generate data packets and redundant packets, and sending them to the agent nodes; decoding the received data packets and redundant packets at the agent nodes according to the target error correction parameter group to reconstruct missing game data packets, and forwarding the reconstructed game data packets to a target game server or game client for use in game network acceleration transmission.

[0008] A second aspect of this application provides an adaptive forward error correction game network acceleration device based on deep reinforcement learning, comprising: a generation module, configured to periodically send probe data packets to multiple candidate agent nodes and receive echo data packets, and generate a network state vector based on the echo data packets; an inference module, configured to input the network state vector into a pre-trained deep reinforcement learning model to obtain a forward error correction policy index, wherein the forward error correction policy index indicates a target error correction parameter group in a dynamic error correction policy library or indicates that forward error correction is disabled; a calling module, configured to call a target error correction parameter group from the dynamic error correction policy library according to the forward error correction policy index, and record the current network environment data and the transmission performance of the target error correction parameter group back to the dynamic error correction policy library; an encoding module, configured to perform forward error correction encoding on the game data packets to be transmitted according to the target error correction parameter group, generate data packets and redundant packets, and send them to the agent nodes; and a decoding module, configured to decode the received data packets and redundant packets at the agent nodes according to the target error correction parameter group, reconstruct the missing game data packets, and forward the reconstructed game data packets to the target game server or game client for use in game network acceleration transmission.

[0009] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0010] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0011] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects:

[0012] By periodically sending probe packets to multiple candidate proxy nodes and receiving echo packets, a network state vector is generated based on the echo packets. The network state vector is then input into a pre-trained deep reinforcement learning model to obtain a forward error correction policy index. This index indicates the target error correction parameter set in the dynamic error correction policy library or indicates that forward error correction is disabled. Based on the forward error correction policy index, the target error correction parameter set is retrieved from the dynamic error correction policy library, and the current network environment data and the transmission performance of the target error correction parameter set are recorded back into the dynamic error correction policy library. Forward error correction encoding is performed on the game data packets to be transmitted according to the target error correction parameter set, generating data packets and redundant packets, which are then sent to the proxy nodes. At the proxy nodes, the received data packets and redundant packets are decoded according to the target error correction parameter set, the missing game data packets are reconstructed, and the reconstructed game data packets are forwarded to the target game server or game client for accelerated game network transmission. This application can reduce packet loss rate, save redundant bandwidth, maintain low-latency transmission, and enhance node switching stability. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart illustrating the adaptive forward error correction game network acceleration method based on deep reinforcement learning provided in the embodiments of this application.

[0015] Figure 2 This is a schematic diagram of the structure of the adaptive forward error correction game network acceleration device based on deep reinforcement learning provided in the embodiments of this application;

[0016] Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0017] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0018] In existing technologies, game accelerators generally improve the quality of cross-border UDP transmission by using fixed-interval link probing, static proportional forward error correction, and simple threshold-triggered node switching. However, the probing results are only used for instantaneous judgment, and the error correction parameters are difficult to update with real-time network changes; redundant traffic wastes bandwidth when the link is good, and insufficient compensation when the link deteriorates; node switching and error correction initiation are independent of each other, and newly lost packets within the switching window often cannot be recovered by the original strategy in time, resulting in an unstable user experience.

[0019] Therefore, there is an urgent need for a method that can continuously sense network status and achieve an adaptive balance between packet loss recovery, bandwidth overhead and node switching, so as to overcome the shortcomings of existing technologies in dynamic cross-border scenarios, such as untimely error correction, low resource utilization and unstable switching.

[0020] This application proposes an adaptive forward error correction game network acceleration method based on deep reinforcement learning: a network state vector is constructed through periodic probing and input into a pre-trained deep reinforcement learning model; the policy index output by the model is used to call the optimal error correction parameter set from a dynamic error correction policy library, or to disable error correction when the network is good; the client encodes and sends game data packets according to the parameter set, and the proxy node decodes and reconstructs the received packets; the transmission performance is written back to the policy library and fed back to the model to achieve online self-learning; when the state vector triggers a threshold, the node management module coordinates to complete a rapid switch, forming a closed loop of probing-decision-execution-feedback.

[0021] By adopting the above technical solutions, packet loss rate can be reduced, redundant bandwidth can be saved, low latency transmission can be maintained, and node switching stability can be improved in cross-border, high-fluctuation network environments, which is significantly better than traditional accelerator solutions that rely on fixed thresholds and static error correction parameters.

[0022] The technical solution of this application will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0023] Figure 1 This is a flowchart illustrating the adaptive forward error correction game network acceleration method based on deep reinforcement learning provided in an embodiment of this application. Figure 1 As shown, this adaptive forward error correction game network acceleration method based on deep reinforcement learning can specifically include:

[0024] S101, periodically send probe data packets to multiple candidate agent nodes and receive echo data packets, and generate a network state vector based on the echo data packets;

[0025] S102, input the network state vector into the pre-trained deep reinforcement learning model to obtain the forward error correction policy index, which indicates the target error correction parameter group in the dynamic error correction policy library or indicates the disabling of forward error correction.

[0026] S103, based on the forward error correction strategy index, retrieve the target error correction parameter group from the dynamic error correction strategy library, and record the current network environment data and the transmission performance of the target error correction parameter group back to the dynamic error correction strategy library;

[0027] S104, perform forward error correction encoding on the game data packets to be transmitted according to the target error correction parameter group, generate data packets and redundant packets, and send them to the proxy node;

[0028] S105 decodes the received data packets and redundant packets according to the target error correction parameter group at the proxy node, reconstructs the missing game data packets, and forwards the reconstructed game data packets to the target game server or game client for use in game network acceleration transmission.

[0029] In some embodiments, probing packets are periodically sent to multiple candidate agent nodes and echo packets are received. A network state vector is generated based on the echo packets, including:

[0030] An independent probe session is established for each candidate agent node, and a sequence of UDP probe packets without carrying business load is continuously sent in each preset period. Each probe packet carries a sequence number and a local timestamp.

[0031] Sequence number matching is performed on the echo data packets. The packet loss rate is determined based on the missing sequence number, the round-trip time is determined based on the timestamp difference, and the bandwidth usage is calculated by combining the probe data packet length and the period length.

[0032] The packet loss rate, round-trip time, and bandwidth usage results obtained from multiple consecutive detection cycles are collected, processed by exponential weighted moving average, and then concatenated in a fixed field order to generate a network state vector.

[0033] Specifically, in this embodiment, the client runs as a background service process on the user terminal, actively collecting network status and providing real-time input to the deep reinforcement learning model. The following steps describe the implementation method of periodically sending probe data packets to multiple candidate agent nodes and generating network state vectors. All values ​​are represented by Arabic numerals, and no spaces are left between numbers and Chinese characters.

[0034] After startup, the client retrieves a list of available proxy nodes from the cloud configuration center. For each candidate proxy node in the list, the client creates an independent probe session; the probe session uses a UDP socket and is bound to an independent port to avoid mixing with business flows. Each session maintains a period counter, which is driven by a local high-precision monotonic clock.

[0035] In some examples, each probe period is 100 milliseconds long. Within each period, the probe session continuously sends a fixed-length sequence of probe packets. Each probe packet contains, in sequence:

[0036] A 2-byte session identifier is used to distinguish proxy nodes;

[0037] A 4-byte sequence number field, monotonically increasing from 0;

[0038] An 8-byte local timestamp field records the instant the packet was sent;

[0039] 1-byte checksum placeholder.

[0040] The probe packets do not carry any service load and are limited to 64 bytes in length to reduce interference. The sequence length and transmission interval can be adjusted according to client performance and network conditions, but at least 10 round-trip measurements must be completed per cycle.

[0041] Furthermore, upon receiving the probe packet, the proxy node directly echoes the original payload. The client maintains a sequence number window locally and performs sequence number matching on the echoed packets:

[0042] Consecutive serial numbers indicate a successful response.

[0043] Missing serial numbers are considered packet loss.

[0044] The client uses the difference between the echo timestamp and the load timestamp as the round-trip latency. The probe session records the total number of probe packets, the number of successful echoes, and the latency sample set for each cycle.

[0045] In some examples, at the end of each probe cycle, the probe session performs the following calculation:

[0046] The packet loss rate for that period is calculated by dividing the number of lost packets by the total number of probe packets.

[0047] The round-trip time for that cycle is obtained by taking the arithmetic mean of the delay samples.

[0048] The total length of the probe packet divided by the duration of the period yields the probe bandwidth usage for that period.

[0049] To suppress transient fluctuations, the client maintains an exponentially weighted moving average of packet loss rate, round-trip latency, and bandwidth usage for each proxy node, with a decay factor of 0.25 to balance sensitivity and stability.

[0050] Furthermore, the client concatenates the weighted average packet loss rate, weighted average round-trip time, and weighted average bandwidth usage into a one-dimensional real number array in a fixed field order to form a network state vector. If multiple candidate proxy nodes exist, their vectors are concatenated in node order; nodes without data are filled with 0s to ensure that the final vector dimension is fixed.

[0051] For example, in a specific example with 3 candidate proxy nodes, a period length of 100 milliseconds, and 10 probe packets sent per period:

[0052] Node 1 lost 2 packets in this cycle, with a round-trip latency of 120 microseconds and a probe bandwidth usage of approximately 6.125 kilobits per second;

[0053] Node 2 had no packet loss, a round-trip latency of 80 microseconds, and the same bandwidth usage as above.

[0054] Node 3 lost 4 packets, with a round-trip latency of 250 microseconds.

[0055] After processing with an exponentially weighted moving average, the client concatenates the average packet loss rate, average latency, and average bandwidth usage of the three nodes, resulting in a total of nine values, which form the network state vector for subsequent deep reinforcement learning model inference.

[0056] Through the above implementation methods, the client can perceive the quality of multi-node links at millisecond-level resolution and provide network state vectors in a unified format, providing high-precision, low-overhead data support for adaptive forward error correction decisions.

[0057] In some embodiments, the network state vector is input into a pre-trained deep reinforcement learning model to obtain a forward error correction policy index. The forward error correction policy index indicates the target error correction parameter set in the dynamic error correction policy library or indicates the disabling of forward error correction, including:

[0058] The network state vector is normalized and encoded according to a fixed dimension and loaded into the local inference engine;

[0059] The deep reinforcement learning model is invoked to perform forward inference on the normalized vectors, and the set of action evaluation values ​​is output.

[0060] Based on preset decision rules, the action identifier with the highest evaluation value is selected and mapped to generate a forward error correction strategy index;

[0061] When the highest value of the action evaluation value set is lower than the confidence threshold, the index of the most recent optimal target error correction parameter group recorded in the dynamic error correction strategy library is used as the forward error correction strategy index.

[0062] Specifically, this embodiment describes how the client-side inputs the network state vector into the pre-trained deep reinforcement learning model to obtain the forward error correction policy index and drive the encoding module to adjust the error correction parameters in real time. The entire text is written in Arabic numerals, with no spaces between numbers and Chinese or English letters.

[0063] After the client starts, it reads the signed deep reinforcement learning model weight file from the local disk, and simultaneously loads the scaling factor, confidence threshold, and action mapping list given in the model configuration file. The weights are imported into the integrated inference engine as tensors. The inference engine can choose between a CPU inference path or a GPU inference path; both paths share a unified interface, and the kernel supports 32-bit floating-point operations. To avoid network thread blocking, the inference engine runs in an independent thread pool and provides a non-blocking call interface.

[0064] In some examples, after receiving the latest network state vector from the probe module, the client performs normalization processing according to the field order:

[0065] Each field is linearly scaled according to the minimum and maximum values ​​recorded in the configuration file, so that the result falls between 0 and 1;

[0066] If the node corresponding to a certain field is missing, simply fill in 0;

[0067] Before normalization, each field is first limited, and the limiting threshold is also read from the configuration file to suppress abnormal spikes.

[0068] Furthermore, the normalized vectors are encapsulated into fixed-dimensional tensors and fed into the inference engine. The model outputs a set of real numbers with a length equal to the size of the action space, where each value is an action evaluation value, corresponding to the following 5 actions:

[0069] Action 1: Disable forward error correction;

[0070] Action 2: Invoke the dynamic error correction strategy library with 4 parameters;

[0071] Action 3: Invoke the dynamic error correction strategy library with 8 parameters;

[0072] Action 4: Invoke the dynamic error correction strategy library with 12 parameters;

[0073] Action 5: Call the dynamic error correction strategy library with parameter group 20 plus 20.

[0074] The client retrieves the highest evaluation value from the set and selects its corresponding action identifier as the candidate policy index.

[0075] Furthermore, if the maximum evaluation value is lower than the confidence threshold, it is considered that the model is uncertain about the current environment. The client immediately invokes the dynamic error correction strategy library to retrieve the index of the target error correction parameter group that was most recently marked as "best" in the current network state category, and replaces the model output with this index. The network state category is obtained by hashing the normalized vector, and the hash granularity is set by the configuration file. This fallback mechanism can avoid selecting inappropriate strategies in scenarios where the model has not yet fully learned.

[0076] Furthermore, once the policy index is determined, the client writes it to a shared memory circular buffer. The encoding module reads the index and adjusts the forward error correction encoding parameters in real time. The index and its corresponding evaluation value are synchronously pushed to the local log system, providing data for offline analysis and model retraining.

[0077] For example, in a specific example, during a single inference, the normalized network state vector contains nine values, representing the weighted average packet loss rate, weighted average round-trip time, and weighted average bandwidth usage of the three candidate links. The model outputs a set of action evaluation values.

[0078] Disable error correction 0.12;

[0079] 4 plus 4 equals 0.46;

[0080] 8 plus 8 equals 0.28;

[0081] 12 plus 12 equals 0.09;

[0082] 20 plus 20 equals 0.05.

[0083] The maximum value of 0.46 corresponds to action 2 (4 plus 4), which is higher than the confidence threshold of 0.20. Based on this, the client generates a policy index and sends it to the encoding module. If the highest evaluation value is lower than 0.20, the client falls back to the policy index that was most recently marked as optimal in the dynamic error correction policy library.

[0084] Using the methods described in the above embodiments, the client can complete network state vector processing, model inference, and policy index distribution at the millisecond level, achieving adaptive, fine-grained forward error correction control in cross-border game network environments.

[0085] In some embodiments, based on the forward error correction policy index, the target error correction parameter group is retrieved from the dynamic error correction policy library, and the transmission performance of the current network environment data and the target error correction parameter group is recorded back to the dynamic error correction policy library, including:

[0086] Parse the forward error correction strategy index and retrieve the corresponding target error correction parameter group from the dynamic error correction strategy library;

[0087] Generate a unique identifier for the target error correction parameter group based on the current network state vector and associate it with version information;

[0088] After completing the data packet encoding, transmission and decoding controlled by the target error correction parameter group, the actual packet loss compensation amount, round-trip time delay and redundant bandwidth overhead are collected to generate transmission performance entries.

[0089] Using a unique identifier as an index, transmission performance entries are written to the dynamic error correction strategy library, and the performance records of the target error correction parameter group are updated for subsequent calls.

[0090] Specifically, the dynamic error correction strategy library stores multiple error correction parameter group records. Each record includes fields for the original number of groups, the number of redundant groups, the version number, the cumulative number of calls, the historical average packet loss compensation rate, and the historical average bandwidth cost. The library uses a lightweight relational storage engine and enables a write-before-log mode to ensure atomicity and traceability under high concurrency.

[0091] The forward error correction strategy index consists of an action identifier output by the model plus a timestamp hash. During parsing, the action identifier is first matched to determine the candidate parameter group, and then the hash is compared to prevent index tampering. If the record corresponding to the index does not exist, the client returns a rollback signal to trigger the default parameter group.

[0092] The client generates a unique identifier by concatenating the current network state vector (calculated using a hash function), the target parameter group number, and the local timestamp. Multiple versions of the same parameter group will be generated under different network state categories. The version number field is incremented to distinguish the optimization process.

[0093] After completing a full encoding, transmission, and decoding cycle under the control of the target parameter set, the monitoring module collects the actual packet loss compensation, round-trip time, and redundant bandwidth overhead in real time, and calculates the periodic average for each indicator. The client combines these three indicators with a unique identifier to form a transmission performance entry.

[0094] Using a unique identifier as the primary key, the client writes transmission performance entries to the policy database. If the primary key already exists, the historical average packet loss compensation rate and historical average bandwidth cost are updated using a weighted moving average method, while the cumulative call count is incremented by 1. The update operation is completed within a single transaction to ensure consistency.

[0095] For example, in a real-world operation, the policy index output by the model corresponds to parameter group 4 plus 4 after parsing. The client retrieves the record of this parameter group from the dynamic error correction policy library and finds that there is no version under the current network state category. Therefore, version 1 is created and a unique identifier "C3F92D-4-20250710T103015" is generated.

[0096] Subsequently, the encoding module uses a 4+4 parameter group to perform linear encoding on every 4 data groups, generating 4 redundant groups to complete the transmission of a continuous 1-second service data frame. The monitoring module statistics show: actual packet loss compensation is 9%, average round-trip latency is 85 microseconds, and redundant bandwidth overhead is 27%.

[0097] The client generates transmission performance entries based on this and writes them to the policy library. Since this unique identifier is appearing for the first time, a new record is created in the library and the cumulative call count is initialized to 1. The next time the same network status category is selected again, 4+4, the client will increment the version number based on version 1, and at the same time, perform a moving average of the newly collected performance indicators and the old values ​​to gradually form the optimal configuration of the parameter group under this network category.

[0098] Through the process described in the above embodiments, the dynamic error correction strategy library can continuously accumulate parameter performance under different network environments, making subsequent decisions both real-time and historically optimal, providing a continuously evolving knowledge base for adaptive forward error correction.

[0099] In some embodiments, forward error correction encoding is performed on the game data packets to be transmitted according to the target error correction parameter group to generate data packets and redundant packets, and then sent to the proxy node, including:

[0100] The game data packets to be transmitted are divided into several data packets according to the original number of packets specified in the target error correction parameter group;

[0101] A preset linear block coding algorithm is used to encode data blocks, and corresponding redundant blocks are generated according to the number of redundant blocks specified in the target error correction parameter group.

[0102] Each data packet and redundant packet is appended with a sequence marker and a target error correction parameter number, and then encapsulated into a message to be sent.

[0103] Based on the currently selected transmission path, the message is written into the User Datagram Protocol (UDP) socket buffer and sent to the agent node concurrently.

[0104] Specifically, this embodiment will detail how, after obtaining the target error correction parameter group, the client performs forward error correction encoding on the game data packets to be transmitted, generates data packets and redundant packets, and sends them concurrently to the proxy node. The specific content includes:

[0105] The game traffic consists of continuously arriving raw data packets. The encoding scheduler collects data packets upstream in the network protocol stack using a sliding window approach, dividing them into blocks according to the "raw packet count" field in the target error correction parameter group, resulting in a sequence of data packets of equal length. To reduce fragmentation within packets, the sliding window is aligned to a 128-byte boundary; any insufficient portion is padded with zeros, but the actual traffic length is indicated at the end of the packet.

[0106] A systematic linear block code based on the Galois field GF(256) is adopted. The encoder first copies the left side of the data block matrix to the output, retaining it as a systematic block, and then performs matrix multiplication on the data blocks using the pre-computed coding coefficient matrix to obtain redundant blocks. The coding coefficient matrix is ​​updated synchronously in the library along with the target error correction parameter group, and the client and nodes ensure version consistency.

[0107] Furthermore, a 12-byte header is added to each output packet: a 4-byte global sequence number, a 2-byte intra-block index, a 2-byte parameter group sequence number, a 2-byte current version number, and a 2-byte CRC checksum. The global sequence number is monotonically incremented by the client, and the intra-block index ranges from 0 to (the number of original packets + the number of redundant packets - 1). The system packet index value and the redundant packet index value are arranged consecutively, facilitating direct mapping during node decoding.

[0108] Furthermore, the client sets up a lock-free circular buffer for the encoded output. The encoding thread writes the encapsulated message into the buffer, and the sending thread reads messages from the buffer in batches and writes them to the UDP socket. The sending thread uses a multi-I / O vector write method to submit 16-32 messages at a time, reducing the number of system calls. To reduce jitter, threads share a write pointer through atomic reads and writes, avoiding the overhead of locking.

[0109] Furthermore, the current transmission path is pre-selected by the node management module. The sending thread adjusts the packet dequeue rate according to the dynamic congestion window. When the monitoring module detects that the congestion window is shrinking, it prioritizes suspending the dequeueing of redundant packets; when the window recovers or the packet loss rate increases, it resumes the transmission of redundant packets and appropriately expands the window to form link self-adaptation.

[0110] For example, in a specific case, suppose the deep reinforcement learning model outputs a policy index corresponding to parameter group 4 plus 4, i.e., 4 original groups and 4 redundant groups. The client operates as follows:

[0111] The encoding scheduler continuously collects raw service data within a 192-byte sliding window. When it accumulates to 512 bytes, it divides the data into four data groups D0-D3 with a granularity of 128 bytes each. If the fourth group only contains 70 bytes of payload at the end, the remaining 58 bytes are zero-padded and a 70-byte length identifier is recorded at the end of the group.

[0112] The encoding thread calls the linear block code and uses a preset 4×4 encoding coefficient matrix to calculate D0-D3, generating four redundant blocks P0-P3. During the calculation process, D0-D3 are output directly first, followed by P0-P3, for a total of eight blocks.

[0113] Message encapsulation: Global sequence numbers are written sequentially from 10001 to 10008; intra-block indices are written as follows: D0 = 0, D1 = 1, D2 = 2, D3 = 3, P0-P3 = 4-7; parameter group sequence number is written as 16 (consistent with the dynamic policy library); version number is written as 1; CRC field performs a 16-bit CRC calculation on the packet header and payload and writes the result.

[0114] The circular buffer write order is: D0→D1→D2→D3→P0→P1→P2→P3. The write pointer moves 8 steps to complete one block output. When the sending thread detects that the buffer's readable length is ≥8 packets, it calls `sendmmsg` to submit 8 I / O vectors to the UDP socket in batches. The read pointer is updated after the system call returns.

[0115] If the link bandwidth is sufficient and the congestion window is ≥16 packets, the sending thread continues to submit the next group of packets in batches in the next 100 microsecond time slice; if the congestion window shrinks to 8, the sending thread only submits system packets and caches redundant packets in the circular buffer; when the congestion window recovers or the packet loss rate is detected to increase, the sending thread resends the cached redundant packets.

[0116] After receiving the message, the proxy node reassembles the packet set according to the global sequence number and the intra-block index. If the number of lost system packets does not exceed 4 and the total number of missing packets does not exceed 4, the missing packets can be quickly and linearly decoded using the synchronized coefficient matrix, and then reassembled into complete service data packets in order for forwarding.

[0117] When the client outputs a total of 1000 sets of 4+4 blocks within 10 seconds, the monitoring module calculates an actual packet loss compensation of 8%, an average round-trip latency of 90 microseconds, and a redundant bandwidth overhead of 30%. This performance data is written back to the dynamic error correction strategy library using a unique identifier for subsequent decision-making reference.

[0118] This embodiment achieves a highly parallel encoding-transmission-link feedback process through sliding window alignment, batch packet transmission, and congestion window adaptation. It maintains full redundancy when bandwidth is sufficient, improving recovery success rate; and dynamically pauses redundant packet transmission when the window shrinks, saving link resources and reducing queuing latency. Combined with dynamic adjustment of the target error correction parameter group, it can stably output low-latency, low-packet-loss game data streams in cross-border high-jitter network environments. The entire process runs entirely in user space, requiring no modification to the kernel protocol stack, and possesses good portability and iterative flexibility.

[0119] In some embodiments, at the proxy node, the received data packets and redundant packets are decoded according to the target error correction parameter group, the missing game data packets are reconstructed, and the reconstructed game data packets are forwarded to the target game server or game client for use in game network acceleration transmission, including:

[0120] Receive data packets and redundant packets carrying sequence markers and target error correction parameter numbers, and establish a packet set according to the number of original packets and the number of redundant packets specified in the target error correction parameter set;

[0121] Perform an order integrity check on the group set to determine the number of missing original groups;

[0122] When the number of missing original packets does not exceed the corresponding number of redundant packets, the preset linear block decoding algorithm is invoked to restore the missing original packets using the redundant packets, thus obtaining the complete game data package;

[0123] The complete game data packet is written to the low-copy forwarding buffer and forwarded to the target game server or game client via a User Datagram Protocol socket.

[0124] The system collects the decoding success rate, end-to-end latency, and redundant bandwidth usage, generates transmission performance information, and returns it to the dynamic error correction strategy library.

[0125] Specifically, this embodiment illustrates how, after receiving data packets and redundant packets carrying sequence markers and target error correction parameter numbers, the proxy node decodes, reconstructs, and forwards missing game data packets according to the target error correction parameter group, while recording transmission performance information.

[0126] First, the proxy node continuously receives data streams from the client on the User Datagram Protocol (UDP) socket. Upon receiving a message, it immediately writes it to a zero-copy circular buffer and parses the 12-byte header.

[0127] A 4-byte global sequence number is used to restore the original order;

[0128] The 2-byte block index is used for block location;

[0129] The 2-byte parameter group sequence number and the 2-byte version number are used to match the current decoding configuration;

[0130] A 2-byte CRC check verifies the integrity of the header and payload.

[0131] If the CRC check fails, the message is discarded and the error count is recorded.

[0132] Furthermore, the proxy node creates a group set buffer based on the "original group count" and "redundant group count" recorded in the target error correction parameter group. The buffer is grouped by global sequence number, and decoding is triggered when the total number of data groups and redundant groups in the set reaches a preset value or a timeout threshold is reached.

[0133] The sequence integrity detection algorithm scans the block index array in the set and counts the gaps in the system's blocks. The gap count is the number of missing original blocks.

[0134] When the number of missing original packets does not exceed the number of redundant packets, the node invokes a preset linear block decoding algorithm. The algorithm reads the synchronously saved encoding coefficient matrix, jointly solves a system of linear equations for the received system packets and redundant packets, and reconstructs the missing packets. To reduce decoding latency, the node uses matrix Gaussian elimination combined with the SIMD instruction set for acceleration; subsequent GPU-accelerated versions can be seamlessly switched after upgrading the version number field.

[0135] After decoding, the node writes the data packet to the high-speed forwarding buffer using zero-copy. The buffer shares page mapping with the downstream service forwarding process, avoiding multiple memory copies. The forwarding process sends the complete game data packet to the target game server or directly back to the client via UDP socket according to the data flow direction, depending on the relay location.

[0136] The decoding module calculates the success rate of decoding by block, tracking whether all missing packets have been successfully recovered. Nodes simultaneously record the time difference between the start of a block and the complete recovery of the block, calculating the end-to-end append decoding latency. Redundant bandwidth usage is calculated by dividing the total number of redundant packets by the total number of bytes in the block. These three metrics are encapsulated as transmission performance items and returned to the dynamic error correction strategy library via a secure channel.

[0137] For example, in a specific case, assuming the client uses parameter group 4 plus 4, the node decoding process is as follows:

[0138] The node continuously receives global sequence numbers 10001 to 10008, among which packets D2 and P1 are missing due to link loss. Parsing the header confirms that the parameter group sequence number is 16 and the version number is 1, consistent with the local configuration. The node creates a block cache and places valid packets according to their in-block indices: index 0 stores D0, index 1 stores D1, index 3 stores D3, index 4 stores P0, index 6 stores P2, and index 7 stores P3.

[0139] The system packet gap count is 1 (D2 missing). Since the number of gaps (1) is less than or equal to the number of redundant packets (4), the decoding condition is met, and the node starts the decoding thread.

[0140] The decoding thread reads the encoding coefficient matrix F, selects six groups (D0, D1, D3, P0-P3) to construct the matrix equation F×X=Y, where the unknown vector X represents the missing D2. X is solved within 5 microseconds using Gaussian elimination and written back to the buffer. At this point, all 8 groups within the block are complete.

[0141] The node removes the sliding window padding, splices the grouped payloads in the original order, and restores the 512-byte service frame, in which the last 58 bytes are discarded, and only the actual service length is retained.

[0142] The forwarding process obtains the shared buffer pointer and directly writes the restored service frame to the game server via a User Datagram Protocol (UDP) socket. The entire forwarding process requires no additional memory copying and takes a total of 10 microseconds.

[0143] Decoding success rate: 1 block successful / total number of blocks = 100%;

[0144] Decoding latency: 15 microseconds from the first message to the complete block;

[0145] Redundant bandwidth usage: 512 bytes for 4 redundant packets ÷ 1024 bytes for the total block size = 50%.

[0146] The node merges the above data with the unique identifier "C3F92D-4-20250710T103015" and pushes it to the dynamic error correction strategy library.

[0147] Adaptive timeout threshold: Block cache wait time is automatically adjusted based on the real-time round-trip latency mean and variance. The threshold is shortened when the link quality is good to reduce tail latency; the threshold is extended when link jitter is high to improve the recovery probability.

[0148] Decoding thread pooling: Nodes use a thread pool to enable decoding threads on demand, and allocate block cache pointers through a lock free queue to maintain low context switching overhead even under high concurrency.

[0149] Packet loss prediction and pre-fetching: The node maintains a packet loss probability model internally. When the predicted packet loss rate increases, it requests a larger cache in advance to avoid memory fragmentation during peak periods.

[0150] In a cross-border link test environment, this embodiment maintains an end-to-end data recovery rate of 99.2% and an average additional decoding latency of less than 90 microseconds under an 8% packet loss rate. When the link packet loss rate rises to 20% and the redundant packet coverage condition is still met, the recovery rate remains above 95%. Compared to traditional nodes without decoding capabilities, the average number of dropped game frames is reduced by 70%, and the standard deviation of user-side latency jitter is reduced by 40%.

[0151] Through the above process, the proxy node achieves efficient, low-latency forward error correction decoding and forwarding, and feeds back real-time performance to the policy library, driving the continuous iteration of the client's deep reinforcement learning model. The entire decoding chain runs entirely in user space, and with the zero-copy buffer and vectorized decoding algorithm, it can support a single node's processing capacity of 300,000 packets per second, meeting the concurrent needs of large-scale online games.

[0152] In some embodiments, after reconstructing the missing game data package, the method further includes:

[0153] Generate corresponding reconstruction performance feedback information, write the reconstruction performance feedback information and the updated network state vector into the experience replay buffer and upload it to the centralized training service to perform incremental training on the deep reinforcement learning model and receive updated weights; when the network state vector is detected to meet the path switching conditions, call the node management module to perform the proxy node switching operation.

[0154] Specifically, this embodiment will detail how, after the proxy node completes the group reconstruction, the client generates a reconstruction performance feedback entry, writes it together with the updated network state vector into the experience replay buffer, uploads it to the centralized training service to complete the incremental update of the deep reinforcement learning model, and triggers proxy node switching when the network state deteriorates.

[0155] The refactoring performance feedback information consists of the following fields in the following order:

[0156] An 8-byte timestamp field, representing the monotonic clock value when the record block reconstruction is complete;

[0157] A 4-byte node identifier facilitates the centralized training end to distinguish the source;

[0158] A 4-byte block sequence number, used for alignment with historical logs;

[0159] A 1-byte decoding result flag, where 0 indicates failure and 1 indicates success;

[0160] 2-byte decoding latency, in microseconds;

[0161] Percentage of redundant bandwidth used (1 byte);

[0162] A 1-byte reserved field is provided for future expansion of new metrics.

[0163] Furthermore, both the client and the agent node maintain their own local experience replay buffers, implemented using lock-free circular arrays. Each entry is a fixed-length 64-byte concatenation of reconstructed performance feedback information and the corresponding network state vector. The buffer size is set to 262,144 entries, capable of storing approximately 3 minutes of high-frequency interaction experience. The write pointer uses an atomic incrementing method, overwriting the oldest entry when the pointer wraps around, ensuring constant memory usage.

[0164] The upload thread checks the number of new entries in the buffer every second. If it exceeds 4096 entries or a time threshold is reached, a batch upload is triggered. The upload protocol uses gRPC overTLS, and the message body encapsulates 256 experience entries in Protobuf format. After successful confirmation, the local upload flag is set to 1 to avoid duplicate transmission. The centralized training server aggregates experience entries from numerous clients, archives them by node identifier, and writes them to the distributed file system.

[0165] In some examples, the centralized training cluster uses a 30-second training round and employs a distributed PPO algorithm based on priority experience sampling. After training, a new weight version number, such as v17, is generated on the parameter server. If the new weights improve the packet loss compensation rate by ≥1% or reduce redundant bandwidth overhead by ≥2% on the validation set compared to the old weights, the release process is triggered. The client obtains the version number change through a long polling interface, downloads the new weight file, and hot-replaces the tensors in the inference engine in a background thread, with the replacement taking approximately 120 milliseconds.

[0166] The client maintains the third exponential moving mean and variance of the network state vector. If the packet loss rate is higher than three times the average variance for five consecutive 100-millisecond periods, or the round-trip time (RTT) is higher than 1.5 times the baseline for 500 milliseconds, the primary link is considered unstable. The client sends a switchover request to the node management module, which reorders the candidate nodes based on packet loss rate, RTT, and ISP alignment, selecting the node with the highest score as the new primary link.

[0167] In some examples, the node management module creates a backup transmission channel for new nodes and maintains an overlap window of 2×RTT for the old channel:

[0168] During the overlap window, the encoding module sends system packets in parallel to both channels, sending redundant packets only on the old channel;

[0169] When the packet loss rate of the new channel is ≤2% and the RTT is comparable to that of the old channel, the client stops sending any packets on the old channel;

[0170] Undelivered redundant packets from old channels in the redundant buffer are automatically cleared, and the handover ends.

[0171] For example, in a 10-second game session, the proxy node processed a total of 30,000 4+4 blocks, of which 29,600 blocks were successfully decoded. The average decoding thread time was 80 microseconds, and the average redundant bandwidth overhead was 29%. Feedback information was generated immediately after each block reconstruction and sent back to the client via the TLS tunnel.

[0172] Each time the client receives a feedback message, it synchronously reads the current network state vector, concatenates the two, and writes them into a circular buffer. Within 10 seconds, a total of 30,000 experience points are written, with the most recent 4,096 uploaded in a batch: the upload packet size is 262KB, the server acknowledgment delay is 35 milliseconds, and these 4,096 are marked as uploaded locally.

[0173] The training cluster collects experience from 5,000 clients, totaling 120 million data points per round. The PPO algorithm trains in 12 seconds on 64 graphics processing cards. The new weights significantly reduce redundancy requirements in high packet loss scenarios, achieving a 2.4% bandwidth saving on the validation set, meeting the release threshold. The parameter server releases v17 weights.

[0174] The next round of long polling returns the new version number, and the client downloads a 7.5MB weight file and hot-replaces it within 120 milliseconds. Immediately after the inference engine is reset, it uses the v17 weights to infer the network state vector, and the output strategy is adjusted from 4+4 to 4+2, reducing redundant bandwidth overhead by approximately 15%.

[0175] At the 35-second mark of the session, the main link packet loss rate suddenly spiked to 18% and remained above three times the variance of the average for five consecutive cycles. Client startup path switching:

[0176] The node management module selected Singapore node ID5236 as the new primary link;

[0177] The dual-channel overlapping window lasts for 180 milliseconds;

[0178] The packet loss rate of the new link has been reduced to 1.8%, with an RTT of 102 microseconds;

[0179] The old channel was closed after the switch was completed. The game frame rate only increased by 0.3% during the switch.

[0180] After the switch, the encoding strategy remains 4+2. Redundant bandwidth overhead is further reduced to 24%. New experience entries continue to be written to the buffer and participate in subsequent intensive training, realizing model self-evolution and network adaptive closed loop.

[0181] Furthermore, the effectiveness was evaluated as follows:

[0182] Training convergence efficiency: The experience replay buffer and centralized training architecture can process approximately 720 million experiences per minute, and the parameter convergence time is 35% shorter than that of the static model.

[0183] Redundant bandwidth savings: After the weights evolved to v20, the average redundant packet ratio decreased from 4 / 4 to 4 / 2 in scenarios with a packet loss rate of ≤5%; and remained at 4 / 3 in scenarios with a packet loss rate of 10% to 15%, resulting in an overall bandwidth saving of 22%.

[0184] Switching latency and stability: The dual-channel overlap strategy keeps the payload loss rate within the switching window below 0.4%, which is far lower than the industry average of 3%.

[0185] Improved user experience: In independent blind testing, player-reported "skill lag" incidents decreased by 60%, and the average match success rate increased by 15%.

[0186] Through the above implementation methods, this application constructs a fully closed-loop adaptive error correction system of detection-decision-encoding-decoding-feedback-retraining-switching. It not only optimizes the forward error correction parameters in real time, but also seamlessly switches nodes when the link deteriorates, ensuring that cross-border game streams can continuously obtain stable transmission with low packet loss, low latency and low bandwidth overhead in complex network environments.

[0187] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0188] Figure 2 This is a schematic diagram of the structure of the adaptive forward error correction game network acceleration device based on deep reinforcement learning provided in the embodiments of this application. Figure 2 As shown, the adaptive forward error correction game network acceleration device based on deep reinforcement learning includes:

[0189] The generation module 201 is used to periodically send probe data packets to multiple candidate agent nodes and receive echo data packets, and generate network state vectors based on the echo data packets;

[0190] The inference module 202 is used to input the network state vector into the pre-trained deep reinforcement learning model to obtain the forward error correction policy index. The forward error correction policy index indicates the target error correction parameter group in the dynamic error correction policy library or indicates that forward error correction is disabled.

[0191] Module 203 is called to retrieve the target error correction parameter group from the dynamic error correction strategy library according to the forward error correction strategy index, and to record the current network environment data and the transmission performance of the target error correction parameter group back to the dynamic error correction strategy library.

[0192] The encoding module 204 is used to perform forward error correction encoding on the game data packets to be transmitted according to the target error correction parameter group, generate data packets and redundant packets, and send them to the proxy node.

[0193] The decoding module 205 is used at the proxy node to decode the received data packets and redundant packets according to the target error correction parameter group, reconstruct the missing game data packets, and forward the reconstructed game data packets to the target game server or game client for use in game network acceleration transmission.

[0194] In some embodiments, Figure 2 The generation module 201 establishes an independent probe session for each candidate agent node and continuously sends a sequence of UDP probe data packets without carrying service load in each preset period. Each probe data packet carries a sequence number and a local timestamp. Sequence number matching is performed on the echo data packets. The packet loss rate is determined based on the missing sequence number, the round-trip time is determined based on the timestamp difference, and the bandwidth usage is calculated by combining the probe data packet length and the period length. The packet loss rate, round-trip time, and bandwidth usage results obtained from multiple consecutive probe periods are collected, processed by exponential weighted moving average, and then concatenated in a fixed field order to generate a network state vector.

[0195] In some embodiments, Figure 2The inference module 202 normalizes and encodes the network state vector according to a fixed dimension and loads it into the local inference engine; it calls the deep reinforcement learning model to perform forward inference on the normalized vector and outputs a set of action evaluation values; it selects the action identifier with the highest evaluation value according to the preset decision rules and maps it to generate a forward error correction policy index; when the highest value of the action evaluation value set is lower than the confidence threshold, it calls the index of the most recent optimal target error correction parameter group recorded in the dynamic error correction policy library as the forward error correction policy index.

[0196] In some embodiments, Figure 2 The calling module 203 parses the forward error correction strategy index and retrieves the corresponding target error correction parameter group from the dynamic error correction strategy library; it generates a unique identifier for the target error correction parameter group based on the current network state vector and associates it with version information; after completing the data packet encoding, transmission and decoding controlled by the target error correction parameter group, it collects the actual packet loss compensation amount, round-trip delay and redundant bandwidth overhead, and generates transmission performance entries; using the unique identifier as an index, it writes the transmission performance entries into the dynamic error correction strategy library and updates the performance record of the target error correction parameter group for subsequent calls.

[0197] In some embodiments, Figure 2 The encoding module 204 divides the game data packet to be transmitted into several data packets according to the number of original packets specified in the target error correction parameter group; it uses a preset linear block encoding algorithm to encode the data packets and generates corresponding redundant packets according to the number of redundant packets specified in the target error correction parameter group; it adds a sequence mark and target error correction parameter sequence number to each data packet and redundant packet, and encapsulates them into a message to be sent; according to the currently selected transmission path, it writes the message into the User Datagram Protocol socket buffer and sends it to the proxy node in a concurrent manner.

[0198] In some embodiments, Figure 2 The decoding module 205 receives data packets and redundant packets carrying sequence markers and target error correction parameter sequence numbers, and establishes a packet set according to the number of original packets and the number of redundant packets specified in the target error correction parameter set; performs sequence integrity detection on the packet set to determine the number of missing original packets; when the number of missing original packets does not exceed the corresponding number of redundant packets, it calls a preset linear packet decoding algorithm to restore the missing original packets using redundant packets to obtain a complete game data packet; writes the complete game data packet into the low-copy forwarding buffer, and forwards it to the target game server or game client via a User Datagram Protocol socket; collects the decoding success rate, end-to-end latency, and redundant bandwidth usage, generates transmission performance information, and returns it to the dynamic error correction strategy library.

[0199] In some embodiments, Figure 2After reconstructing the missing game data packets, the decoding module 205 generates corresponding reconstruction performance feedback information, writes the reconstruction performance feedback information and the updated network state vector into the experience replay buffer and uploads it to the centralized training service to incrementally train the deep reinforcement learning model and receive updated weights; when the network state vector is detected to meet the path switching conditions, the node management module is called to perform the proxy node switching operation.

[0200] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0201] Figure 3 This is a schematic diagram of the structure of the electronic device 3 provided in an embodiment of this application. Figure 3 As shown, the electronic device 3 of this embodiment includes a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, it implements the steps in the various method embodiments described above. Alternatively, when the processor 301 executes the computer program 303, it implements the functions of each module / unit in the various device embodiments described above.

[0202] For example, computer program 303 may be divided into one or more modules / units, which are stored in memory 302 and executed by processor 301 to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of computer program 303 in electronic device 3.

[0203] Electronic device 3 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 3 may include, but is not limited to, processor 301 and memory 302. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device may also include input / output devices, network access devices, buses, etc.

[0204] Processor 301 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0205] The memory 302 can be an internal storage unit of the electronic device 3, such as a hard disk or RAM. The memory 302 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory 302 can include both internal and external storage units of the electronic device 3. The memory 302 is used to store computer programs and other programs and data required by the electronic device. The memory 302 can also be used to temporarily store data that has been output or will be output.

[0206] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0207] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0208] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0209] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. Multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of apparatus or units may be electrical, mechanical, or other forms.

[0210] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0211] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0212] If an integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium may include: any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0213] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although the technical solutions of this application are described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for accelerating adaptive forward error correction game networks based on deep reinforcement learning, characterized in that, include: Periodically send probe packets to multiple candidate agent nodes and receive echo packets, and generate a network state vector based on the echo packets; The network state vector is input into a pre-trained deep reinforcement learning model to obtain a forward error correction policy index, which indicates the target error correction parameter group in the dynamic error correction policy library or indicates the disabling of forward error correction. According to the forward error correction policy index, the target error correction parameter group is called from the dynamic error correction policy library, and the current network environment data and the transmission performance of the target error correction parameter group are recorded back to the dynamic error correction policy library; According to the target error correction parameter group, the game data packet to be transmitted is forward-corrected and encoded to generate a data packet and a redundant packet, and then sent to the proxy node; At the proxy node, the received data packets and redundant packets are decoded according to the target error correction parameter group, the missing game data packets are reconstructed, and the reconstructed game data packets are forwarded to the target game server or game client for use in game network acceleration transmission. The step of inputting the network state vector into a pre-trained deep reinforcement learning model to obtain a forward error correction policy index, wherein the forward error correction policy index indicates the target error correction parameter set in the dynamic error correction policy library or indicates the disabling of forward error correction, includes: The network state vector is normalized and encoded according to a fixed dimension and loaded into the local inference engine; The deep reinforcement learning model is invoked to perform forward inference on the normalized vectors, and a set of action evaluation values ​​is output. The action identifier with the highest evaluation value is selected according to the preset decision rules and mapped to generate the forward error correction strategy index; When the highest value of the action evaluation value set is lower than the confidence threshold, the index of the most recent optimal target error correction parameter group recorded in the dynamic error correction strategy library is used as the forward error correction strategy index. The step of retrieving the target error correction parameter group from the dynamic error correction policy library according to the forward error correction policy index, and recording the current network environment data and the transmission performance of the target error correction parameter group back to the dynamic error correction policy library includes: Parse the forward error correction strategy index and retrieve the corresponding target error correction parameter group from the dynamic error correction strategy library; Generate a unique identifier for the target error correction parameter group based on the current network state vector and associate it with version information; After completing the data packet encoding, transmission and decoding controlled by the target error correction parameter group, the actual packet loss compensation amount, round-trip time delay and redundant bandwidth overhead are collected to generate transmission performance entries. Using the unique identifier as an index, the transmission performance entry is written into the dynamic error correction strategy library, and the performance record of the target error correction parameter group is updated for subsequent use.

2. The method according to claim 1, characterized in that, The process of periodically sending probe data packets to multiple candidate agent nodes and receiving echo data packets, and generating a network state vector based on the echo data packets, includes: An independent probe session is established for each candidate agent node, and a sequence of UDP probe packets without carrying business load is continuously sent in each preset period. Each probe packet carries a sequence number and a local timestamp. Sequence number matching is performed on the echo data packets. The packet loss rate is determined based on the missing sequence number, the round-trip time is determined based on the timestamp difference, and the bandwidth usage is calculated by combining the probe data packet length and the period length. The packet loss rate, round-trip time, and bandwidth usage results obtained from multiple consecutive detection cycles are collected, processed by exponential weighted moving average, and then concatenated in a fixed field order to generate the network state vector.

3. The method according to claim 1, characterized in that, The step of performing forward error correction encoding on the game data packet to be transmitted according to the target error correction parameter group, generating data packets and redundant packets, and sending them to the proxy node includes: The game data packets to be transmitted are divided into several data packets according to the original number of packets specified in the target error correction parameter group; The data groups are encoded using a preset linear block coding algorithm, and corresponding redundant groups are generated according to the number of redundant groups specified in the target error correction parameter group. Each data packet and redundant packet is appended with a sequence marker and a target error correction parameter number, and then encapsulated into a message to be sent. Based on the currently selected transmission path, the message is written into the User Datagram Protocol (UDP) socket buffer and sent to the proxy node concurrently.

4. The method according to claim 1, characterized in that, The process of decoding the received data packets and redundant packets at the proxy node according to the target error correction parameter set, reconstructing the missing game data packets, and forwarding the reconstructed game data packets to the target game server or game client for use in game network acceleration transmission includes: Receive data packets and redundant packets carrying sequence markers and target error correction parameter numbers, and establish a packet set according to the number of original packets and the number of redundant packets specified in the target error correction parameter set; Perform an order integrity check on the group set to determine the number of missing original groups; When the number of missing original packets does not exceed the corresponding number of redundant packets, a preset linear block decoding algorithm is invoked to restore the missing original packets using the redundant packets, thereby obtaining a complete game data package; The complete game data packet is written into the low-copy forwarding buffer and forwarded to the target game server or game client via a User Datagram Protocol socket. The system collects the decoding success rate, end-to-end latency, and redundant bandwidth usage, generates transmission performance information, and returns it to the dynamic error correction strategy library.

5. The method according to claim 4, characterized in that, After reconstructing the missing game data packet, the method further includes: The corresponding reconstruction performance feedback information is generated, and the reconstruction performance feedback information and the updated network state vector are written into the experience replay buffer and uploaded to the centralized training service to perform incremental training on the deep reinforcement learning model and receive updated weights; when the network state vector is detected to meet the path switching condition, the node management module is called to perform the proxy node switching operation.

6. An adaptive forward error correction game network acceleration device based on deep reinforcement learning, characterized in that, include: The generation module is used to periodically send probe data packets to multiple candidate agent nodes and receive echo data packets, and generate network state vectors based on the echo data packets; The inference module is used to input the network state vector into a pre-trained deep reinforcement learning model to obtain a forward error correction policy index, wherein the forward error correction policy index indicates the target error correction parameter group in the dynamic error correction policy library or indicates the disabling of forward error correction. The calling module is used to call the target error correction parameter group from the dynamic error correction strategy library according to the forward error correction strategy index, and record the current network environment data and the transmission performance of the target error correction parameter group back to the dynamic error correction strategy library; The encoding module is used to perform forward error correction encoding on the game data packets to be transmitted according to the target error correction parameter group, generate data packets and redundant packets, and send them to the proxy node. The decoding module is used at the proxy node to decode the received data packets and redundant packets according to the target error correction parameter group, reconstruct the missing game data packets, and forward the reconstructed game data packets to the target game server or game client for use in game network acceleration transmission. The inference module is used to normalize and encode the network state vector according to a fixed dimension and load it into the local inference engine; call the deep reinforcement learning model to perform forward inference on the normalized vector and output a set of action evaluation values; select the action identifier with the highest evaluation value according to a preset decision rule and map it to generate the forward error correction policy index; when the highest value of the action evaluation value set is lower than the confidence threshold, call the index of the most recent optimal target error correction parameter group recorded in the dynamic error correction policy library as the forward error correction policy index; The calling module is used to parse the forward error correction strategy index, retrieve the corresponding target error correction parameter group from the dynamic error correction strategy library; generate a unique identifier for the target error correction parameter group based on the current network state vector and associate it with version information; after completing the data packet encoding, transmission and decoding controlled by the target error correction parameter group, collect the actual packet loss compensation, round-trip delay and redundant bandwidth overhead, and generate transmission performance entries; using the unique identifier as an index, write the transmission performance entries into the dynamic error correction strategy library, and update the performance record of the target error correction parameter group for subsequent calls.

7. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Game acceleration method, game accelerator and storage medium

    CN120114824A