Information data transmission method based on computer network
By deploying intelligent agents at both the sending and receiving ends, and employing collaborative decision-making and network coding techniques, the network congestion and fairness issues caused by cross-cloud data migration tools were resolved, achieving stable and efficient data transmission.
Patent Information
- Application Number
- CN202511196142.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-12-02
AI Technical Summary
Existing cross-cloud data migration tools, in their pursuit of efficiency, tend to create "elephant flows," leading to network congestion and fairness imbalances, and making it difficult to avoid operator QoS intervention.
Intelligent agents are deployed at both the sending and receiving ends to collaboratively acquire network state parameters, form a state view of the transmission path, and make collaborative decisions based on the policy network to determine the target transmission rate. By combining network coding and rate shaping techniques, the transmission strategy is dynamically adjusted to optimize network traffic.
It achieves fairness and stability in network transmission, reduces the impact on other service flows, avoids network congestion and operator QoS intervention, and improves transmission efficiency and reliability.
Smart Images

Figure CN121056533A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer network and data communication technology, and in particular to an information data transmission method based on a computer network. Background Technology
[0002] In the current enterprise IT architecture, multi-cloud deployment of business systems and data storage has become the mainstream model; the resulting cross-cloud data migration needs are becoming increasingly frequent, often involving data volumes of hundreds of terabytes; such migration tasks usually rely on dedicated tools provided by cloud service providers or self-built channels based on open-source technology stacks to achieve data movement under specific network conditions.
[0003] To cope with large-scale data migration, most existing transmission solutions employ techniques such as multi-threaded parallelism, data segmentation, and incremental synchronization. Some existing technologies introduce intelligent rate adjustment mechanisms, which can dynamically adjust the sending window based on the processing capabilities of terminal nodes. These methods aim to improve the throughput of a single migration task and strive to complete the data transfer within a predetermined time.
[0004] However, when such data flows traverse the operator's backbone network, their sustained high bandwidth consumption can easily form "elephant flows." This type of traffic can occupy network queue resources for a long time, potentially causing localized network congestion and affecting the transmission quality of other data flows at the same time. The optimization goals of existing transmission mechanisms are mainly aimed at task completion efficiency, and in cross-domain environments, it is difficult to avoid the intervention of operator QoS policies, which may also cause periodic fluctuations in network status. Summary of the Invention
[0005] In view of the aforementioned existing problems, the present invention is proposed.
[0006] This invention provides a computer network-based information data transmission method to address the problems of existing cross-cloud data migration tools, which, in pursuit of efficiency, tend to create massive data flows, leading to network congestion, fairness imbalances, and difficulty in avoiding operator QoS intervention.
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0008] This invention provides a method for transmitting information data based on a computer network, applied to a data transmission system including a sending node and a receiving node, characterized by comprising the following steps:
[0009] Step S1: Deploy a sending agent on the sending node and a receiving agent on the receiving node.
[0010] Step S2: The transmitting end agent and the receiving end agent cooperate to acquire network status parameters and form a status view of the transmission path within the control cycle;
[0011] Step S3: Based on the state view, the transmitting agent and the receiving agent make collaborative decisions and determine the target transmission rate for the current control cycle.
[0012] In step S4, the sending node divides the data to be transmitted into blocks according to the target transmission rate, and sends it to the receiving node through the transmission path in a packet rate shaping and throttling manner, whereby the receiving node reassembles and restores the original data.
[0013] In a preferred embodiment of the information data transmission method based on a computer network according to the present invention, the network status parameters include at least one of the following:
[0014] The receiving agent periodically measures and reports end-to-end latency or its stability indicators, packet loss rate, or explicit congestion flags, in this case; and the sending agent observes actual throughput or queuing delay trends.
[0015] The receiving agent sends network status parameters back to the sending agent via feedback messages, which can be sent independently or carried in reverse service data.
[0016] As a preferred embodiment of the information data transmission method based on computer networks described in this invention, the collaborative decision-making includes:
[0017] The sending agent proposes a rate suggestion value based on its local state and sends it to the receiving agent.
[0018] The receiving agent evaluates the proposed rate value based on its local state and returns a confirmed rate value.
[0019] The transmitting agent uses the rate confirmation value as the target transmission rate and re-updates it in the next control cycle based on the latest status.
[0020] As a preferred embodiment of the information data transmission method based on computer network described in this invention, the target transmission rate is obtained by the online rate adjustment action given by the policy network, and the policy network updates with the network state parameters as input and the goal of maximizing the cumulative expected revenue.
[0021] The benefit is determined by a reward function, and the reward function at least simultaneously reflects the trade-off between throughput improvement, latency stability improvement, and rate change smoothness.
[0022] The reward function is defined as follows:
[0023] In each control cycle k, the transmitting and receiving ends obtain: instantaneous throughput, stability index of round-trip delay sequence, and rate change amplitude between adjacent actions, respectively; direction alignment and [0, 1] interval normalization are performed on each index to obtain z. T,k z D,k z S,k Normalization employs piecewise linearity and truncation.
[0024] The three types of normalized scores are represented as follows: Throughput improvement term z T,k The larger the value, the better; the latency stability improvement term z D,k The larger the value, the better; rate smoothing term z S,k The bigger the better;
[0025] In period k, an immediate reward is given by a weighted sum and used to accumulate the target return of the policy network:
[0026] r k =ω T,k z T,k +ω D,k z D,k +ω S,k z S,k ,
[0027] Where, r k This represents the instantaneous reward for period k, in dimensionless units, where k is the control period index, and z... T,k z D,k z S,k These represent the normalized scores for throughput improvement, latency stability improvement, and rate smoothness, respectively, with values ranging from [0, 1], ω. T,k ω D,k ω S,k This represents the corresponding weight, which is non-negative and the sum of the three is 1;
[0028] The weights are obtained from static priors and online self-tuning; static priors The online phase adapts and updates based on differences in goal achievement:
[0029]
[0030] u T,k =α T (1-z T,k ), u D,k =α D (1-z D,k ), u S,k =α S (1-z S,k ),
[0031] Where, ω j,k This represents the weight of indicator j in period k, where j takes the values {T, D, S}, uj,k ε represents the weight driving force. w Let α be the numerical stability constant. T α D α S α is the weighted sensitivity coefficient; when the receiver reports packet loss or the explicit congestion flag increases, α D The limit will be temporarily raised in the next cycle to curb unstable actions;
[0032] For r k Perform an exponential moving average to reduce short-term jitter, and set upper and lower bounds to clip to [-1, 1];
[0033] The smoothed r k For policy gradients or actor-critic feedback channels, the advantage estimation and baseline implementation uses a built-in caliber that synchronously resets the advantage estimation window length when weight updates change significantly.
[0034] As a preferred embodiment of the information data transmission method based on computer networks described in this invention, the policy network adopts policy gradient or actor-commentator update based on the advantage function, and introduces an exploration term or entropy regularization term in the update to achieve stable convergence.
[0035] The policy network update rules include:
[0036] Collect (s) within the control period k a k r k s k+1 The sequence is arranged into trajectory batches in chronological order. If a replay buffer is used, only the commentator branch is sampled using the off-policy, while the actor branch uses the near-on-policy batch.
[0037] Calculate discount returns and construct advantages using a sliding window. Critic's branch output V φ (s k Using this as a baseline, the advantage is to perform zero-mean and unit-variance normalization;
[0038] In each update round, the joint loss is minimized to simultaneously drive the actors and critics:
[0039]
[0040] in, Represents joint loss, π θ (a k |s k Let θ represent the conditional probability of the actor's action given by the parameter θ. V represents the advantage estimate. φ (s kThis represents the state value given by the commentator using the parameter φ. λ represents the experience discount return. V λ represents the weight of the value item. H Represents the entropy regularity coefficient. E represents the policy distribution entropy, used to encourage moderate exploration. k The expectation is the mean of step k within the current training batch;
[0041] In the formula, the generalized dominance estimation is used to construct...
[0042]
[0043] δ k =r k +γ R V φ (s k+1 )-V φ (s k ),
[0044] Where, γ R λ represents the return discount factor. A δ represents the dominance attenuation factor. k L represents the timing difference error. k This indicates the cutoff length starting from step k, in steps.
[0045] As a preferred embodiment of the information data transmission method based on computer networks described in this invention, the sending node performs packet-level throttling on data packets according to the target transmission rate, ensures that the average transmission rate is lower than the target transmission rate in any control period through a token bucket or equivalent rate shaper, and allocates bandwidth budgets according to weights among multiple concurrent connections.
[0046] As a preferred embodiment of the information data transmission method based on computer network described in this invention, the block transmission includes network encoding of the data to be transmitted to generate multiple encoded data blocks, so that the receiving node can still decode and restore the data even if some parts are lost.
[0047] The network coding is preferably fountain code or an equivalent on-demand redundancy coding method.
[0048] As a preferred embodiment of the information data transmission method based on computer networks described in this invention, the multiple concurrent connections adopt a weighted round-robin or equivalent scheduling strategy to control the transmission order, and adaptively adjust the bandwidth weight of each connection within the control period based on the latency stability index fed back by the receiver, so as to reduce the queue expansion caused by long-term occupation of a single connection.
[0049] As a preferred embodiment of the information data transmission method based on computer networks described in this invention, the receiving end intelligent agent carries a congestion indication in the feedback message, including a delay transition alarm, a sudden increase in packet loss alarm, or an explicit congestion marker exceeding the threshold alarm. Based on this, the sending end intelligent agent enters a conservative mode and alleviates path congestion by reducing the target transmission rate or increasing the coding redundancy.
[0050] As a preferred embodiment of the information data transmission method based on computer networks described in this invention, the sending node and the receiving node are cloud computing virtual machine instances or object storage service endpoints belonging to different cloud service providers, and the method is implemented through an application layer protocol or a control channel running above the transport layer.
[0051] The beneficial effects of this invention are:
[0052] This invention intelligently maintains the fairness and stability of network transmission. Through collaborative perception and decision-making between the sending and receiving agents, this method no longer pursues the maximization of speed for a single data stream in isolation, but considers its own behavior within the global perspective of the entire network environment. Its core algorithm proactively generates pattern-friendly and smoothly fluctuating network traffic by balancing multiple objectives such as throughput, latency stability, and rate smoothness. This significantly reduces the impact on other service flows in the shared network link, avoids congestion caused by instantaneous queue accumulation, fundamentally promotes fair bandwidth sharing among different data flows, and effectively suppresses the negative interference of operator QoS policies.
[0053] The intelligent decision engine embedded in this invention can perceive subtle changes in network status in real time and dynamically adjust the focus of the transmission strategy through an adaptive mechanism of reward function weights. For example, when early signs of congestion such as increased packet loss or latency spikes are detected, the system can automatically prioritize ensuring transmission stability and quickly and smoothly reduce speed to prevent congestion from worsening. Combined with the application of network coding technology, even in scenarios where packet loss is unavoidable, the receiving end can still efficiently decode and recover data, significantly reducing the additional latency and signaling overhead caused by retransmission requests. This ensures overall transmission efficiency and success rate in complex, variable, and unreliable cross-carrier network environments. Attached Figure Description
[0054] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation on the scope of this application.
[0055] Figure 1 This is a flowchart illustrating the information data transmission method based on a computer network in this embodiment. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0057] All terms used in this application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0058] The terms “first”, “second”, etc., used in this application are merely used to distinguish similar objects and differentiate the first object from another, and are not used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance.
[0059] This application proposes an information data transmission method based on computer networks, combined with... Figure 1 As shown, the method includes:
[0060] Step S1: Deploy a sending agent on the sending node and a receiving agent on the receiving node.
[0061] In this embodiment, the sending agent and the receiving agent refer to program units running in user space on their respective nodes, possessing interfaces for state monitoring, collaborative decision-making, and rate execution. They are deployed in parallel with the existing transmission protocol stack without modifying the kernel. Specifically, they read throughput, round-trip latency, and packet loss statistics through a local acquisition interface, and exchange summarized state and decision results through a control channel. The control cycle is 100ms by default and can be adjusted from 50 to 500ms, selected based on the round-trip latency distribution and jitter handling of the cross-domain link. The initial handshake timeout is 2 control cycles by default and can be adjusted to [1, 5] control cycles, determined by offline simulation and engineering experience. If either agent is unavailable during the handshake, it can optionally degenerate to single-end monitoring and use a conservative rate limit until the other end recovers and completes capability negotiation. If necessary, if local resources are insufficient (CPU or memory usage exceeds a preset threshold), this embodiment suspends policy updates but maintains the previous effective target transmission rate unchanged.
[0062] Step S2: The sending agent and the receiving agent work together to acquire network state parameters and form a state view of the transmission path within the control cycle.
[0063] Step S3: Based on the state view, the sending agent and the receiving agent make collaborative decisions and determine the target transmission rate for the current control cycle.
[0064] Step S4: The sending node divides the data to be transmitted into blocks according to the target transmission rate, and sends them to the receiving node through the transmission path in a packet rate shaping and throttling manner, and the receiving node reassembles and restores the original data.
[0065] In this embodiment, the block size is 256 KiB by default, adjustable [64 KiB, 4 MiB], determined by a combination of storage I / O efficiency and link MTU fragmentation risk; packet-level rate shaping is performed using token bucket, with the average rate aligned to the target sending rate. The burst limit is set by default to the amount that can be sent within one round-trip time (adjustable to 0.5-2 times), and is adjusted in each control cycle based on the latest acknowledgment value; for example, sending is scheduled at millisecond-level clock speeds to reduce short-term queue jitter. Optionally, if the underlying library provides equivalent leaky bucket / precise pacing capabilities, they are directly reused and input / output consistency is maintained. If necessary, when the local queue backlog exceeds a threshold or disk writes are limited, the upper limit of the target sending rate is temporarily reduced to prevent backpressure propagation.
[0066] In one embodiment, the network state parameters include at least one of the following:
[0067] The receiving agent periodically measures and reports end-to-end latency or its stability indicators, packet loss rate, or explicit congestion flags, in this case; and the sending agent observes actual throughput or queuing delay trends.
[0068] The receiving agent sends parameters back to the sending agent via feedback messages. These feedback messages can be sent independently or carried in reverse business data.
[0069] Furthermore, the feedback message includes a cycle number, a relative timestamp, end-to-end latency statistics (either mean or jitter), packet loss count or explicit congestion flag (in this example), and an optional queuing trend indicator. In this embodiment, the measurement window is 500ms by default and can be adjusted from 200-1000ms, based on the requirement that the window covers multiple control cycles and can smooth short-term spikes. For example, latency stability is calculated using the sliding standard deviation of the round-trip latency within the window, which is then normalized by piecewise linearity and truncation as the reported component. Similarly, packet loss and explicit congestion flags (in this example) are truncated using quantiles to suppress outliers. The feedback message is sent at the same frequency as the control cycle by default. If the reverse service flow is reused, it is carried with additional metadata without affecting the service. Optionally, when the reverse path is unstable, it switches to an independent control channel and reduces the sending frequency to once every two control cycles to reduce overhead. In case of missing measurements, the sending end uses the most recent valid feedback and local observation interpolation to complete one control cycle. If there are more than three consecutive missing measurements, it enters conservative mode.
[0070] In one embodiment, collaborative decision-making includes:
[0071] The sending agent proposes a rate suggestion value based on its local state and sends it to the receiving agent.
[0072] The receiving agent evaluates the proposed rate value based on its local state and returns a confirmed rate value.
[0073] The sending agent uses the rate confirmation value as the target transmission rate and re-updates it in the next control cycle based on the latest status. Specifically, to avoid rate oscillations, the change in the target transmission rate is limited within adjacent control cycles: the minimum step size defaults to 2% of the current rate, and the maximum step size defaults to no more than 20%, adjustable [10%, 30%], based on the trade-off between delay inflation and completion time in the experiment. When negotiation times out (no confirmation received), the sending agent lowers the candidate rate to 90% of the previously confirmed rate and retryes negotiation in the next control cycle. Similarly, when the difference between the suggested values of both parties exceeds a preset deviation threshold (default 15%), the receiver's evaluation is used as the standard, and a conflict event is recorded for subsequent weight adjustment. In abnormal situations (three consecutive conflicts or two consecutive congestion indications), this embodiment freezes the rate increase, maintains it, or slightly decreases it until the stable index recovers above the set threshold.
[0074] In one embodiment, the target transmission rate is obtained by the policy network giving online rate adjustment actions. The policy network takes network state parameters as input and updates the rate with the goal of maximizing the cumulative expected revenue.
[0075] The revenue is determined by the reward function, and the reward function at least simultaneously reflects the trade-off between throughput improvement, latency stability improvement and rate change smoothness.
[0076] The reward function is defined as follows:
[0077] In each control cycle k, the transmitting and receiving ends obtain: instantaneous throughput, stability index of round-trip delay sequence (such as the sliding standard deviation of jitter), and rate change amplitude between adjacent actions, respectively; direction alignment and [0,1] interval normalization are performed on each index to obtain z. T,k z D,k z S,k Normalization uses piecewise linear and truncation methods, with a default window size of 500ms, adjustable to [200, 1000]ms.
[0078] The three types of normalized scores are represented as follows: Throughput improvement term z T A larger k is better, and the latency stability improvement term z D,k The larger the value, the better (it increases as delay jitter decreases), the rate smoothing term z. S,k The larger the better (increases as the change in motion decreases);
[0079] In period k, an immediate reward is given by a weighted sum and used to accumulate the target return of the policy network:
[0080] r k =ω T,k z T,k +ω D,k z D,k +ω S,k z S,k ,
[0081] Where, r k This represents the instantaneous reward for period k, in dimensionless units, where k is the control period index, and z... T,k z D,k z S,k These represent the normalized scores for throughput improvement, latency stability improvement, and rate smoothness, respectively, with values ranging from [0, 1], ω. T,k ω D,k ω S,k This represents the corresponding weight, which is non-negative and the sum of the three is 1;
[0082] The weights are obtained from static priors and online self-tuning; static priors The online phase adapts and updates based on differences in goal achievement:
[0083]
[0084] u T,k =α T (1-z T,k ), u D,k =α D (1-z D,k ), u S,k =α S (1-z s,k ),
[0085] Where, ω j,k This represents the weight of indicator j in period k, where j takes the values {T, D, S}, u j,k ε represents the weight driving force. w This is the numerical stability constant, with a default value of 10. -3 α T α D α S α is the weighted sensitivity coefficient, default (1.0, 1.0, 0.6), adjustable [0.2, 2.0]; when the receiver reports packet loss or the explicit congestion flag increases, α... D The limit will be temporarily raised in the next cycle to curb unstable actions;
[0086] For r kAn exponential moving average is used to reduce short-term jitter. The smoothing coefficient is 0.6 by default and can be adjusted to [0.3, 0.9]. The upper and lower bounds are set to be clipped to [-1, 1] to avoid extreme values dominating the learning process.
[0087] The smoothed r k For policy gradients or actor-critic feedback channels, the advantage estimation and baseline implementation adopts a built-in caliber, and the advantage estimation window length is reset synchronously when the weight update changes significantly.
[0088] Specifically, the throughput boost component in the reward comes from the effective throughput gain within the measurement window, the latency stability component is mapped from the decrease in round-trip latency jitter, and the rate smoothing component is mapped from the decrease in the amplitude of changes in adjacent actions. In this embodiment, the three components undergo direction consistency and interval normalization within the same window. Normalization employs piecewise linearity and truncation to limit the impact of outliers. Furthermore, the online self-adjustment of weights is driven by the degree of non-achievement, and the attention of stability weights is temporarily increased when congestion signals occur. The reward is subjected to exponential moving average and pruning to improve training stability. Its smoothing and pruning caliber are decoupled from the control cycle, and the window length is set in engineering based on the link fluctuation cycle. If necessary, if the reward is abnormal (such as missing tests or sudden distortion), the current update is skipped and the previous effective feedback is used.
[0089] Specifically, in the reward function, throughput, latency stability, and rate smoothing are placed on the same comparable scale, and weights that adapt to changes in the scenario are given. First, normalized scores for the three categories are obtained using a unified window, and then weighted summation is used to form an immediate reward, taking into account both transmission efficiency and interaction smoothness. The weight part adopts a lightweight mechanism of prior and online self-adjustment, using the gap in achieving the target as the driving quantity to obtain the normalized weight allocation, and temporarily increasing the focus on stability when congestion signals occur. This avoids the rigidity of fixed weights under diverse networks and does not introduce complex multi-objective outer layer optimization, making it easy to implement in real-time control cycles. Smoothing and pruning are used to suppress learning instability caused by transient anomalies, and the interface layer remains decoupled from the existing policy update module, reducing the intrusion into the original training pipeline.
[0090] In one embodiment, the policy network employs a policy gradient or actor-commentator update based on the advantage function, and introduces an exploration term or entropy regularization term into the update to ensure stable convergence;
[0091] Policy network update rules include:
[0092] Collect (s) within the control period k a k r k s k+1The sequence is arranged into trajectory batches in chronological order. If a replay buffer is used, only the commentator branch is sampled using the off-policy method, while the actor branch uses the near-on-policy batch method. The batch length is 2048 by default and can be adjusted to [512, 4096].
[0093] Calculate discount returns and construct advantages using a sliding window. Critic's branch output V φ (s k Using this as a baseline, it is advantageous to perform zero-mean and unit-variance normalization to improve numerical stability;
[0094] In each update round, the joint loss is minimized to simultaneously drive the actors and critics:
[0095]
[0096] in, Represents joint loss, π θ (a k |s k Let θ represent the conditional probability of the actor's action given by the parameter θ. V represents the advantage estimate. φ (s k This represents the state value given by the commentator using the parameter φ. λ represents the experience discount return. V Indicates the weight of the value item, defaults to 0.5, adjustable [0.1, 2.0], λ H This represents the entropy regularization coefficient, with a default value of 0.01 and an adjustable value of [0, 0.05]. E represents the policy distribution entropy, used to encourage moderate exploration. k The expectation is the mean of step k within the current training batch;
[0097] In the formula, the generalized dominance estimation is used to construct...
[0098]
[0099] δ k =r k +γ R V φ (s k+1 )-V φ (s k ),
[0100] Where, γ R This represents the return discount factor, defaulting to 0.99, adjustable [0.90, 0.999], λ. A This represents the advantage attenuation factor, defaulting to 0.95, adjustable [0.80, 0.99], δ k L represents the timing difference error.k This indicates the cutoff length starting from step k, in steps.
[0101] The stabilization terms and update details are as follows:
[0102] (a) KL Gating: Online monitoring of the gap between old and new strategies KL If the mean exceeds the threshold The default value is 0.02, adjustable from [0.01, 0.10]. This round stops early and the learning rate is reduced. KL This represents the average KL divergence over the current batch;
[0103] (b) Gradient norm clipping: The gradient of each branch is denoted by g... max Global norm clipping, default 0.5, adjustable [0.3, 1.0], g max Indicates the upper limit of the clipping;
[0104] (c) Learning rate strategy: η lr Default 3×10 -4 Supports linear preheating and cosine annealing, η lr Indicates the base learning rate;
[0105] (d) Critics' Target Smoothing: Critics employ a target network φ tgt Soft update coefficient τ V Default is 0.995, adjustable [0.90, 0.999], φ tgt τ represents the target commentator parameter. V This is a soft update coefficient;
[0106] (e) Small batches and rounds: Each round divides the batch into N mb (The number of samples in each mini-batch is 1024, adjustable [256, 4096]) and iterate E. ep Number of rounds (default 4, adjustable [1, 10]), N mb The number of samples in each mini-batch;
[0107] After several rounds of parameter updates, θ is fixed, and the target transmission rate action is given in the next control cycle. When d KL When there is a drastic change in value loss or entropy reduction rate, shorten the update interval and record the rollback checkpoint;
[0108] In this embodiment, the observation vector of the policy network consists of a summary of network state parameters from the most recent control cycles, including throughput, latency stability, packet loss and explicit congestion marking, and bandwidth weight adjustment records. The action space is the adjustment amount of the target transmission rate, which is adjusted upwards or downwards based on the previous confirmed rate by a limit, and is constrained by a safety upper and lower limit. Optionally, to reduce policy drift, if the policy change rate monitored online reaches a preset threshold, short-term convergence is performed with a smaller learning step size and a more frequent evaluation interval, and if necessary, the parameters are rolled back to the most recent checkpoint. In abnormal situations, if the observation vector is incomplete or significantly distorted, the action remains unchanged within that cycle and is marked as an invalid sample.
[0109] Specifically, the three objectives are coupled together using joint loss: the strategy term is driven by advantage, the value term reduces the estimation bias of returns through mean squared error, and the entropy term maintains the necessary exploration intensity; the advantage estimation adopts a generalized form, which controls the trade-off between short-term and long-term information through the two parameters of discount and decay, and is further combined with zero-mean standardization to mitigate variance.
[0110] In terms of stabilization, KL gating is used as an early stopping signal to limit policy drift, gradient norm pruning suppresses abnormal gradients in a single step, learning rate scheduling handles the different sensitivities in the early and late stages of training, and target commentator introduces time smoothing to reduce the jitter of the guiding signal; mini-batch and multi-round are improved multiple times on the same data to improve sample utilization; after the default values and adjustable range of the whole process are given, it is easy to quickly implement and port under different link conditions, while reserving an interface for subsequent integration with the reward weight adaptive mechanism;
[0111] In one embodiment, the sending node performs packet-level throttling on data packets according to the target sending rate, ensures that the average sending rate is lower than the target sending rate in any control period through a token bucket or equivalent rate shaper, and allocates bandwidth budgets according to weights among multiple concurrent connections.
[0112] In one embodiment, chunked transmission includes network encoding of the data to be transmitted to generate multiple encoded data blocks, so that the receiving node can still decode and restore the data even if some parts are lost.
[0113] Network coding is preferably fountain coding or an equivalent on-demand redundancy coding method;
[0114] Furthermore, the granularity of the encoded symbols is 1-4 KiB by default, and the redundancy is initially set to 5%-15% of the original data block count and can be dynamically adjusted according to link packet loss. The receiver stops requesting subsequent encoded blocks when it reaches the recoverable threshold (the original data block count plus the minimum redundancy tolerance) to reduce invalid transmissions. Optionally, when the feedback message indicates a short-term increase in packet loss rate but the latency stability is still within the threshold, the encoding redundancy is preferentially increased rather than immediately reduced to maintain throughput continuity; when the redundancy has reached its upper limit and still cannot meet the decoding progress, this embodiment synchronously triggers a conservative rate reduction.
[0115] In one embodiment, multiple concurrent connections use a weighted round-robin or equivalent scheduling strategy to control the sending order, and adaptively adjust the bandwidth weight of each connection within the control period based on the latency stability index fed back by the receiver, so as to reduce the queue expansion caused by long-term occupation of a single connection.
[0116] In this embodiment, the number of concurrent connections is 2-8 by default. The bandwidth weights are normalized and summed to 1. The upper limit of the weight for a single connection is no more than 0.8 and the lower limit is no less than 0.05 by default. The weight change rate between adjacent control cycles is limited to within 20% to avoid oscillation. Specifically, if the latency stability of the flow corresponding to a connection deteriorates and is accompanied by a congestion indication, its weight is reduced by a preset attenuation factor in the next control cycle, and the released quota is allocated to the remaining connections in the same manner. Similarly, when all connections are stable, the weights are gradually restored to the initial evenly distributed configuration. In abnormal situations (short-term connection inactivity or continuous error codes), the weight of the connection is temporarily reduced to the lower limit and then recovers in a slow-start manner after recovery.
[0117] In one embodiment, the receiving agent carries a congestion indication in the feedback message, including a delay transition alarm, a sudden increase in packet loss alarm, or an explicit congestion mark ratio exceeding the threshold alarm. Based on this, the sending agent enters a conservative mode and alleviates path congestion by reducing the target transmission rate or increasing coding redundancy.
[0118] In this embodiment, the triggering criteria for a delay transition alarm are that the average round-trip delay within a measurement window increases by more than a set threshold compared to its historical steady-state average, with a default threshold of 30-50ms. A packet loss surge alarm is triggered when the packet loss rate exceeds 0.5% within a window or increases by more than three times compared to the previous window. An explicit congestion marking ratio exceeding the threshold is triggered by a default 2%. Upon entering conservative mode, the target transmission rate is immediately reduced proportionally, with a default reduction of 10%-30%, and a cooldown period of 2-3 control cycles is set to avoid frequent fluctuations. Optionally, if the congestion signal is only manifested in packet loss rather than delay transition, the coding redundancy is first increased to the upper limit before a slight rate reduction. The exit condition is that all indicators recover to within the threshold for several consecutive windows without any new alarms.
[0119] In one embodiment, the sending node and the receiving node are cloud computing virtual machine instances or object storage service endpoints belonging to different cloud service providers, and the method is implemented through an application layer protocol or a control channel running above the transport layer, without relying on the proprietary functions of the underlying network devices.
[0120] In this embodiment, the control channel reuses the application layer load carrying status and decision metadata of the reverse service connection by default, and can be adjusted to an independent transport layer session to be compatible with scenarios without reverse services. To adapt to the address translation and security policies of different cloud environments, the heartbeat and keep-alive interval is 30 seconds by default, and can be adjusted from 10 to 60 seconds, which is determined by the cloud-side idle reclamation policy and NAT maintenance rules. When the intermediate box discards control metadata on the cross-provider path, this embodiment switches to a more fault-tolerant lightweight encoding bearer and reduces the frequency of control messages to improve the penetration rate, keeping the input and output of the method consistent without relying on the underlying proprietary functions.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0122] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of this application and form different embodiments. For example, all the embodiments above can be used in any combination. The information disclosed in this background section is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.
Claims
1. A method for transmitting information data based on a computer network, applied to a data transmission system including a sending node and a receiving node, characterized in that, Includes the following steps: Step S1: Deploy a sending agent on the sending node and a receiving agent on the receiving node. Step S2: The transmitting end agent and the receiving end agent cooperate to acquire network status parameters and form a status view of the transmission path within the control cycle; Step S3: Based on the state view, the transmitting agent and the receiving agent make collaborative decisions and determine the target transmission rate for the current control cycle. In step S4, the sending node divides the data to be transmitted into blocks according to the target transmission rate, and sends it to the receiving node through the transmission path in a packet rate shaping and throttling manner, whereby the receiving node reassembles and restores the original data.
2. The information data transmission method based on a computer network as described in claim 1, characterized in that, The network state parameters include at least one of the following: The receiving agent periodically measures and reports end-to-end latency or its stability indicators, packet loss rate, or explicit congestion flags, in this case; and the sending agent observes actual throughput or queuing delay trends. The receiving agent sends network status parameters back to the sending agent via feedback messages, which can be sent independently or carried in reverse service data.
3. The information data transmission method based on a computer network as described in claim 1, characterized in that, The collaborative decision-making includes: The sending agent proposes a rate suggestion value based on its local state and sends it to the receiving agent. The receiving agent evaluates the proposed rate value based on its local state and returns a confirmed rate value. The transmitting agent uses the rate confirmation value as the target transmission rate and re-updates it in the next control cycle based on the latest status.
4. The information data transmission method based on a computer network as described in claim 1, characterized in that, The target transmission rate is obtained by the policy network giving online rate adjustment actions. The policy network takes the network state parameters as input and updates it with the goal of maximizing the cumulative expected revenue. The benefit is determined by a reward function, and the reward function at least simultaneously reflects the trade-off between throughput improvement, latency stability improvement, and rate change smoothness. The reward function is defined as follows: In each control cycle k, the transmitting and receiving ends obtain: instantaneous throughput, stability index of round-trip delay sequence, and rate change amplitude between adjacent actions, respectively; direction alignment and [0, 1] interval normalization are performed on each index to obtain z. T,k z D,k z S,k Normalization employs piecewise linearity and truncation. The three types of normalized scores are represented as follows: Throughput improvement term z T,k The larger the value, the better; the latency stability improvement term z D,k The larger the value, the better; rate smoothing term z S,k The bigger the better; In period k, an immediate reward is given by a weighted sum and used to accumulate the target return of the policy network: r k =ω T,k z T,k +ω D,k z D,k +ω S,k z S,k , Where, r k This represents the instantaneous reward for period k, in dimensionless units, where k is the control period index, and z... T,k z D,k z S,k These represent the normalized scores for throughput improvement, latency stability improvement, and rate smoothness, respectively, with values ranging from [0, 1], ω. T,k ω D,k ω S,k This represents the corresponding weight, which is non-negative and the sum of the three is 1; The weights are obtained from static priors and online self-tuning; static priors The online phase adapts and updates based on differences in goal achievement: at T,k =α T (1-z T,k ),at D,k =α D (1-z D,k ),at S,k =α S (1-z S,k ), Where, ω j,k This represents the weight of indicator j in period k, where j takes the values {T, D, S}, u j,k ε represents the weight driving force. w Let α be the numerical stability constant. T α D α S α is the weighted sensitivity coefficient; when the receiver reports packet loss or the explicit congestion flag increases, α... D The limit will be temporarily raised to the upper limit in the next cycle to curb unstable actions; For r k Perform an exponential moving average to reduce short-term jitter, and set upper and lower bounds to clip to [-1, 1]; The smoothed r k For policy gradients or actor-critic feedback channels, the advantage estimation and baseline implementation uses a built-in caliber that synchronously resets the advantage estimation window length when weight updates change significantly.
5. The information data transmission method based on a computer network as described in claim 4, characterized in that, The policy network employs policy gradient or actor-commentator updates based on the advantage function, and introduces an exploration term or entropy regularization term in the update to ensure stable convergence. The policy network update rules include: Collect (s) within the control period k a k ,r k s k+1 The sequence is arranged into trajectory batches in chronological order. If a replay buffer is used, only the commentator branch is sampled using the off-policy, while the actor branch uses the near-on-policy batch. Calculate discount returns and construct advantages using a sliding window. Critic's branch output V φ (s k Using this as a baseline, the advantage is to perform zero-mean and unit-variance normalization; In each update round, the joint loss is minimized to simultaneously drive the actors and critics: in, Represents joint loss, π θ (a k |s k Let θ represent the conditional probability of the actor's action given by the parameter θ. V represents the advantage estimate. φ (s k This represents the state value given by the commentator using the parameter φ. λ represents the experience discount return. V λ represents the weight of the value item. H Represents the entropy regularity coefficient. E represents the policy distribution entropy, used to encourage moderate exploration. k The expectation is the mean of step k within the current training batch; In the formula, the generalized dominance estimation is used to construct... δ k =r k +γ R V φ (s k+1 )-V φ (s k ), Where, γ R λ represents the return discount factor. A δ represents the dominance attenuation factor. k L represents the timing difference error. k This indicates the cutoff length starting from step k, in steps.
6. The information data transmission method based on a computer network as described in claim 1, characterized in that, The sending node performs packet-level throttling on data packets according to the target sending rate, and ensures that the average sending rate is lower than the target sending rate in any control period through a token bucket or equivalent rate shaper, and allocates bandwidth budgets according to weights among multiple concurrent connections.
7. The information data transmission method based on a computer network as described in claim 1, characterized in that, The block transmission includes network encoding of the data to be transmitted to generate multiple encoded data blocks, so that the receiving node can still decode and restore the data even if some parts are lost. The network coding is preferably fountain code or an equivalent on-demand redundancy coding method.
8. The information data transmission method based on a computer network as described in claim 6, characterized in that, The multiple concurrent connections are controlled by a weighted round-robin or equivalent scheduling strategy. Based on the latency stability index fed back by the receiver, the bandwidth weight of each connection is adaptively adjusted within the control period to reduce queue expansion caused by long-term occupation of a single connection.
9. The information data transmission method based on a computer network as described in claim 2, characterized in that, The receiving agent carries congestion indications in the feedback message, including delay transition alarms, packet loss surge alarms, or explicit congestion marking alarms exceeding the threshold. Based on this, the sending agent enters a conservative mode, which alleviates path congestion by reducing the target transmission rate or increasing coding redundancy.
10. The information data transmission method based on a computer network as described in claim 1, characterized in that, The sending node and the receiving node are cloud computing virtual machine instances or object storage service endpoints belonging to different cloud service providers, and the method is implemented through an application layer protocol or a control channel running above the transport layer.