Network data transmission method for route optimization based on deep reinforcement learning and related equipment

By employing a deep reinforcement learning-based routing optimization method that combines routing and caching strategies, this approach addresses the challenge of achieving multi-QoS optimization in complex networks using traditional methods. It improves network resource utilization and transmission efficiency, enabling highly efficient and reliable data transmission.

CN121887703APending Publication Date: 2026-04-17FIBRLINK NETWORKS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FIBRLINK NETWORKS
Filing Date
2025-11-28
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional routing optimization and data transmission mechanisms struggle to achieve comprehensive optimization of multiple service quality indicators in complex, large-scale converged networks. They are unable to effectively cope with sudden changes in network status and fluctuations in service load, resulting in low transmission efficiency and uneven resource allocation.

Method used

A deep reinforcement learning-based approach is adopted to obtain network environment state information, jointly optimize routing and caching strategies, generate a joint strategy that includes link weights, key cache nodes and cache data stream types, and combine the association information of the data to be transmitted to determine a transmission scheme that minimizes transmission time under preset constraints.

Benefits of technology

It improves network resource utilization, reduces data transmission latency, and achieves efficient and reliable network data transmission. It can also achieve collaborative optimization of multiple QoS indicators in complex converged network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121887703A_ABST
    Figure CN121887703A_ABST
Patent Text Reader

Abstract

The invention provides a network data transmission method for route optimization based on deep reinforcement learning and related equipment. The method comprises the following steps: acquiring environment state information of a network environment; joint optimization of a routing strategy and a caching strategy is carried out based on the environment state information, and joint strategy information about routing and caching is obtained; the joint strategy information comprises a link weight, a key cache node and a cache data flow type; acquiring transmission associated information of the to-be-transmitted data; based on the transmission association information, determining data transmission information with minimum transmission duration when a preset constraint is met; and executing a routing action based on the joint policy information, executing a data scheduling action based on the data transmission information, and transmitting the to-be-transmitted data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of communications, and in particular to a network data transmission method and related equipment based on deep reinforcement learning for routing optimization. Background Technology

[0002] With the deepening of the convergence of telecommunications, broadcasting, and the internet, modern communication networks are evolving towards heterogeneity, service diversification, and dynamism, posing severe challenges to traditional routing optimization and data transmission mechanisms. In complex, large-scale converged networks, network topology changes dynamically, and service types are diverse and heterogeneous. Traditional routing algorithms based on static rules or single indicators struggle to achieve comprehensive optimization of multiple Quality of Service (QoS) indicators (such as latency, bandwidth, and packet loss rate), and are also unable to effectively cope with sudden changes in network state and fluctuations in service load, leading to increasingly prominent problems such as low transmission efficiency and uneven resource allocation. Summary of the Invention

[0003] In view of this, the purpose of this disclosure is to propose a network data transmission method and related equipment based on deep reinforcement learning for routing optimization.

[0004] In a first aspect, this disclosure provides a network data transmission method for routing optimization based on deep reinforcement learning, comprising: Obtain environmental status information of the network environment, wherein the environmental status information is used to indicate link utilization information and data flow information of the network environment; Based on the environmental state information, a joint optimization of routing and caching strategies is performed to obtain joint strategy information for routing and caching; the joint strategy information includes link weights, key cache nodes, and cache data stream types; Obtain transmission association information of the data to be transmitted, wherein the transmission association information is used to indicate the data information, data link information and data scheduling information of the data to be transmitted; Based on the transmission association information, determine the data transmission information that minimizes the transmission time while satisfying preset constraints; Based on the joint policy information, routing actions are performed, and based on the data transmission information, data scheduling actions are performed to transmit the data to be transmitted.

[0005] A second aspect of this disclosure provides a network data transmission apparatus for routing optimization based on deep reinforcement learning, comprising: An environment acquisition module is used to acquire environmental status information of the network environment, wherein the environmental status information is used to indicate link utilization information and data flow information of the network environment; The joint optimization module is used to jointly optimize routing and caching strategies based on the environmental state information to obtain joint strategy information about routing and caching; the joint strategy information includes link weights, key cache nodes, and cache data stream types. The transmission information acquisition module is used to acquire transmission association information of the data to be transmitted, wherein the transmission association information is used to indicate the data information, data link information and data scheduling information of the data to be transmitted. The transmission optimization module is used to determine, based on the transmission association information, the data transmission information that minimizes the transmission time while satisfying preset constraints; The action execution module is used to execute routing actions based on the joint policy information and data scheduling actions based on the data transmission information to transmit the data to be transmitted.

[0006] A third aspect of this disclosure provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in the first aspect.

[0007] In a fourth aspect, this disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method described in the first aspect.

[0008] A fifth aspect of this disclosure provides a computer program product including computer program instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect.

[0009] As described above, this disclosure provides a network data transmission method and related equipment based on deep reinforcement learning for route optimization. By acquiring network environment state information (including link utilization and data flow information), a joint optimization strategy is generated, comprising link weights, key cache nodes, and cached data flow types. Simultaneously, combining the associated information of the data to be transmitted (data, link, and scheduling information), a transmission scheme minimizing transmission time is determined under preset constraints. Finally, routing selection is performed based on the joint strategy, and data is scheduled according to the transmission scheme. This effectively improves network resource utilization, reduces data transmission latency, and achieves efficient and reliable network data transmission. This disclosure provides an integrated method for route optimization and data transmission based on deep reinforcement learning. It can achieve collaborative optimization of multiple QoS indicators through a single intelligent agent model and coordinate different optimization objectives based on a reward function mechanism, ultimately achieving efficient, stable, and adaptive data transmission in complex and integrated network environments. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of a network data transmission architecture based on deep reinforcement learning for routing optimization, according to an embodiment of this disclosure.

[0012] Figure 2 This is a schematic diagram of the structure of an exemplary electronic device according to an embodiment of the present disclosure.

[0013] Figure 3 This is a schematic flowchart illustrating a network data transmission method for routing optimization based on deep reinforcement learning, according to an embodiment of this disclosure.

[0014] Figure 4 This is a schematic diagram of routing modeling based on deep reinforcement learning, according to an embodiment of this disclosure.

[0015] Figure 5 This is a flowchart of a routing optimization algorithm based on deep reinforcement learning, according to an embodiment of this disclosure.

[0016] Figure 6 This is a flowchart of a network data transmission algorithm based on deep reinforcement learning, according to an embodiment of this disclosure.

[0017] Figure 7 This is a schematic diagram of a network data transmission device for routing optimization based on deep reinforcement learning, according to an embodiment of this disclosure. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0020] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0021] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0022] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0023] Figure 1 A schematic diagram of a network data transmission architecture based on deep reinforcement learning for routing optimization, according to an embodiment of this disclosure, is shown. (Reference) Figure 1 The network data transmission architecture 100 based on deep reinforcement learning for routing optimization may include a server 110, a terminal 120, and a network 130 providing a communication link. The server 110 and the terminal 120 can be connected via a wired or wireless network 130. The server 110 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, security services, and CDN.

[0024] Terminal 120 can be implemented in hardware or software. For example, when terminal 120 is implemented in hardware, it can be various electronic devices with a display screen and support page display, including but not limited to smartphones, tablets, e-book readers, laptops, and desktop computers. When terminal 120 is implemented in software, it can be installed in the electronic devices listed above; it can be implemented as multiple software programs or software modules (e.g., software programs or software modules used to provide distributed services) or as a single software program or software module, without specific limitations.

[0025] It should be noted that the network data transmission method based on deep reinforcement learning for routing optimization provided in this embodiment can be executed by either the terminal 120 or the server 110. It should be understood that... Figure 1 The number of terminals, networks, and servers shown is for illustrative purposes only and is not intended to be a limitation. Any number of terminals, networks, and servers can be used depending on implementation needs.

[0026] Figure 2 A schematic diagram of the hardware structure of an exemplary electronic device 200 provided in an embodiment of this disclosure is shown. For example... Figure 2 As shown, the electronic device 200 may include: a processor 202, a memory 204, a network module 206, a peripheral interface 208, and a bus 210. The processor 202, memory 204, network module 206, and peripheral interface 208 are interconnected within the electronic device 200 via the bus 210.

[0027] Processor 202 may be a central processing unit (CPU), neural network processor (NPU), microcontroller (MCU), programmable logic device, digital signal processor (DSP), application-specific integrated circuit (ASIC), or one or more integrated circuits. Processor 202 can be used to perform functions related to the techniques described in this disclosure. In some embodiments, processor 202 may also include multiple processors integrated as a single logic component. For example, such as... Figure 2 As shown, processor 202 may include multiple processors 202a, 202b and 202c.

[0028] Memory 204 can be configured to store data (e.g., instructions, computer code, etc.). Figure 2As shown, the data stored in memory 204 may include program instructions (e.g., program instructions for implementing the network data transmission method for routing optimization based on deep reinforcement learning according to embodiments of this disclosure) and data to be processed (e.g., the memory may store configuration files of other modules, etc.). Processor 202 may also access the program instructions and data stored in memory 204 and execute the program instructions to operate on the data to be processed. Memory 204 may include volatile storage devices or non-volatile storage devices. In some embodiments, memory 204 may include random access memory (RAM), read-only memory (ROM), optical disk, magnetic disk, hard disk, solid-state drive (SSD), flash memory, memory stick, etc.

[0029] Network module 206 can be configured to provide communication with other external devices to electronic device 200 via a network. This network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, Near Field Communication (NFC), etc.), a cellular network, the Internet, or a combination thereof. It is understood that the type of network is not limited to the specific examples described above. In some embodiments, network module 206 may include any combination of any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, etc.

[0030] The peripheral interface 208 can be configured to connect the electronic device 200 to one or more peripheral devices to enable information input and output. For example, peripheral devices may include input devices such as keyboards, mice, touchpads, touch screens, microphones, and various sensors, as well as output devices such as displays, speakers, vibrators, and indicator lights.

[0031] Bus 210 can be configured to transmit information between various components of electronic device 200 (e.g., processor 202, memory 204, network module 206, and peripheral interface 208), such as internal buses (e.g., processor-memory bus), external buses (USB port, PCI-E bus), etc.

[0032] It should be noted that although the architecture of the above-described electronic device 200 only shows the processor 202, memory 204, network module 206, peripheral interface 208, and bus 210, in specific implementations, the architecture of the electronic device 200 may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the architecture of the above-described electronic device 200 may only include the components necessary for implementing the embodiments of this disclosure, and does not necessarily include all the components shown in the figures.

[0033] In related technologies, traditional methods for route optimization typically employ heuristic algorithms or centralized control strategies, lacking real-time awareness of the network environment and autonomous decision-making capabilities. Although some research has attempted to introduce machine learning methods into route decision-making, most models remain limited to a single optimization objective or rely on multi-model ensembles, resulting in high computational overhead and difficulty in achieving cross-domain, multi-objective collaborative optimization in real-world networks. Furthermore, existing routing mechanisms generally lack the ability to dynamically reconstruct content distribution paths, failing to achieve joint optimization of caching and routing based on real-time network conditions and business needs, severely limiting content delivery efficiency.

[0034] In terms of network data transmission, traditional transmission algorithms are mostly based on fixed strategies or threshold control, lacking the ability to adapt to dynamic changes in link load. Especially in converged scenarios with multiple concurrent services, different services have significantly different requirements for transmission latency and reliability, and existing methods struggle to achieve differentiated data transmission scheduling, leading to a decline in the experience of high-priority services. Furthermore, traditional methods are inadequate in load balancing, easily causing network congestion and uneven resource utilization, further impacting overall transmission performance.

[0035] In recent years, deep reinforcement learning (DRL) technology has shown significant potential in the field of network optimization. Through continuous interaction between agents and the environment, it can achieve autonomous decision-making and policy optimization in unknown dynamic environments. Existing research has attempted to apply DRL to single aspects of routing or transmission control, such as value function-based routing optimization or congestion avoidance mechanisms, but a unified framework for multi-objective joint optimization and coordinated routing and transmission control is still lacking. In particular, in terms of reward function design, existing methods often focus on a single optimization objective (such as minimizing latency or maximizing throughput), lacking systematic modeling of the trade-offs between multiple QoS indicators, leading to conflicting optimization objectives or unstable convergence performance.

[0036] In view of this, this disclosure provides a network data transmission method and related equipment based on deep reinforcement learning for routing optimization. By acquiring network environment state information (including link utilization and data flow information), a joint optimization strategy is generated, including link weights, key cache nodes, and cached data flow types. Simultaneously, combining the associated information of the data to be transmitted (data, link, and scheduling information), a transmission scheme minimizing transmission time is determined under preset constraints. Finally, routing selection is performed based on the joint strategy, and data is scheduled according to the transmission scheme. This effectively improves network resource utilization, reduces data transmission latency, and achieves efficient and reliable network data transmission. This disclosure provides an integrated method for routing optimization and data transmission based on deep reinforcement learning, which can achieve collaborative optimization of multiple QoS indicators through a single intelligent agent model and coordinate different optimization objectives based on a reward function mechanism, ultimately achieving efficient, stable, and adaptive data transmission in complex and integrated network environments.

[0037] See Figure 3 , Figure 3 A schematic flowchart illustrating a network data transmission method for route optimization based on deep reinforcement learning according to an embodiment of the present disclosure is shown. This network data transmission method for route optimization based on deep reinforcement learning according to an embodiment of the present disclosure can be deployed on a server. Figure 3 In the network data transmission method 300 based on deep reinforcement learning for routing optimization, the following steps may be further included.

[0038] In step S310, environmental status information of the network environment is obtained. The environmental status information is used to indicate the link utilization information and data flow information of the network environment.

[0039] This can be achieved by deploying various monitoring tools and sensors within the network to comprehensively collect link utilization information, such as real-time bandwidth utilization and transmission latency of each link; simultaneously, data flow information can be acquired, covering the source address, destination address, traffic volume, and transmission direction of the data flow. Environmental status information allows for the determination of the network's current operational status. Link utilization information helps identify congested and idle links in the network, preventing data flows from passing through congested links and causing increased transmission latency; data flow information reveals the characteristics and requirements of different data flows, thus enabling the development of personalized routing strategies for different types of data flows. Ultimately, based on this information, more scientific and reasonable routing planning can be achieved, effectively improving the overall transmission efficiency of the network, reducing data transmission latency and packet loss rate, and enhancing network stability and reliability.

[0040] In step S320, the routing strategy and caching strategy are jointly optimized based on the environmental state information to obtain joint strategy information about routing and caching; the joint strategy information includes link weight, key cache nodes, and cache data stream type.

[0041] Link weight, in the context of network topology, measures the importance or transmission capacity of a link in data transmission. Link weight is calculated based on performance parameters such as bandwidth, latency, and packet loss rate. A higher weight indicates better link performance, and routing will tend to select links with higher weights. Key caching nodes refer to nodes with significant caching capabilities within the network. These nodes may be located in critical network positions, such as core routers and edge access nodes. Key caching nodes can store large amounts of frequently used data or content. When other nodes request this data, they can retrieve it directly from the key caching nodes, thereby reducing data transmission latency and network bandwidth consumption. Cached data stream types refer to the types or characteristics of data that need to be cached in the network. Different data stream types have different access patterns and importance, such as static web pages, video streams, and audio streams. Different caching strategies can be developed based on the characteristics of the data stream types to improve caching efficiency and hit rate.

[0042] The routing and caching strategies are jointly optimized based on environmental state information (such as network topology, link performance parameters, and user request patterns). First, by monitoring and analyzing the network environment in real time, various performance indicators of the links are obtained, and the weight of each link is calculated to reflect its transmission capacity and reliability. Simultaneously, based on the frequency and distribution of user requests, key cache nodes in the network are identified; these nodes should have high storage capacity and good network connectivity. Then, corresponding caching strategies are formulated for different types of cached data streams. For example, prefetching and cache replacement strategies can be used for popular video streams to improve cache hit rate. Finally, the routing and caching strategies are combined to generate joint policy information on routing and caching, including link weights, key cache nodes, and cached data stream types, guiding data transmission and cache management in the network. This improves overall network performance and user experience. Specifically, the reasonable allocation of link weights allows data transmission to select the optimal path, reducing transmission latency and packet loss rate, and improving network reliability and stability. The identification and utilization of key cache nodes enable the rapid retrieval of frequently used data, reducing the load on the source server and decreasing the amount of data transmitted over the network, thus saving network bandwidth resources. Caching strategies tailored to different types of cached data streams improve the targeting and effectiveness of caching, further enhancing cache hit rates and allowing users to obtain the content they need faster. This effectively optimizes network resource usage, improves network transmission efficiency and response speed, and provides users with a smoother and more efficient network service.

[0043] In some embodiments, the environmental state information includes: ; in, For at any time Environmental status information, For at any time Link The load, For at any time Link capacity, for The weight, for The weight, For at any time The hardware readings for the real-time queue usage of node m's port. For at any time The port line speed of node m is negotiated with the physical layer. Based on the aforementioned environmental state information, a joint optimization of routing and caching strategies is performed to obtain joint strategy information regarding routing and caching, including: Based on the trained joint optimization network, feature extraction is performed on the environmental state information to obtain environmental state features; and based on the environmental state features, the link weights of links and key cache nodes in the network environment are determined. The corresponding cache data stream type is determined based on the hardware type of the key cache node; Based on the link weight, the key cache node, and the cache data stream type, the joint policy information at time u is determined; wherein, the joint policy information includes: ;in, For the joint policy information at time u, For the link at time u The link weights, where F is the set of links. For the critical cache node at time u, At time u, in the cache node The set of nodes at time u For the key cache node at time u The cached data stream type.

[0044] Among them, by constructing a pre-trained joint optimization network, the time step is optimized. uThis solution employs deep feature extraction and analysis of environmental status information to dynamically determine link weight allocation and key cache node selection. First, it calculates the weight of each link based on environmental characteristics such as link load and capacity ratio, and node queue status, identifying key nodes with high caching value. Then, it matches the optimal cache data stream type (e.g., video stream, real-time interactive data) based on node hardware characteristics. Finally, it generates a joint optimization decision including link weight configuration, key cache node deployment, and corresponding caching strategies. This solution achieves intelligent collaboration between routing and cache management. By dynamically adapting to changes in link load and node hardware capabilities, it significantly improves network resource utilization and reduces transmission latency. Simultaneously, the accurate matching of key node caching strategies effectively improves cache hit rate and data acquisition efficiency, comprehensively optimizing data transmission performance and user experience in complex network environments.

[0045] In some embodiments, method 300 further includes: The participant network and the commenter network are trained based on historical policy learning records, and the trained participant network is determined as the joint optimization network; wherein, Historical learning trajectory information of a preset length is sampled from the historical strategy learning records, and features are extracted from the historical learning trajectory information to obtain historical learning trajectory features; The joint reward function and the loss function between the participant network and the commentator network are determined based on the historical learning trajectory features. The network parameters of the initial joint network are updated based on the joint reward function to minimize the loss function, thereby obtaining the trained participant network. The joint reward function includes: Among them, the end-to-end latency reward value , For at any time Maximum end-to-end latency of all data streams For at any time End-to-end latency; packet loss rate reward value , For at any time Packet loss rate; Load balancing reward value , For at any time Load balancing coefficient; hop count bonus , For at any time The number of hops in the path. This represents the maximum number of hops across all possible paths in the network environment. The weights for end-to-end latency reward values. The weight of the packet loss rate reward value, The weights for load balancing reward values. The weight of the route hop count reward value; The link weights and nodes that maximize the joint reward function are determined as the link weights and the critical cache nodes.

[0046] This study employs a reinforcement learning framework to collaboratively train a participant-commentator network using historical policy learning records, thereby constructing an intelligent joint optimization network. First, trajectory information is sampled and features extracted from historical data. A joint reward function is designed, incorporating four dimensions: end-to-end latency, packet loss rate, load balancing, and hop count (each metric is normalized and configured with adjustable weights). The commentator network evaluates the policy's effectiveness and calculates the loss function, guiding the participant network to maximize rewards through parameter updates. Finally, the well-trained participant network serves as the joint optimization network. Its output link weight configuration and key cache node selection strategy comprehensively balance transmission efficiency (reducing latency and hop count), stability (reducing packet loss), and resource utilization (optimizing load balancing). By automatically uncovering network performance optimization patterns through multi-objective reinforcement learning, this approach effectively overcomes the limitations of traditional heuristic algorithms, enabling adaptive policy adjustments in dynamic network environments and significantly improving data transmission reliability, real-time performance, and resource utilization efficiency.

[0047] Specifically, such as Figure 4 As shown, when performing route optimization, it is necessary to ensure that the designed scheme is compatible with existing Internet Protocol (IP) networks. To address this issue, this invention reallocates the weights of all links at the beginning of each time step and derives the routing strategy based on the reconstructed link weights. a. Routing scheme: time The routing scheme can be represented as: .

[0048] in, Indicates at time The reassigned link weights, where F represents the link set. The link weights are actually mapped to the forwarding table entries and interface metrics of the router / switch. After the weights are distributed, they take effect at hardware line speed in the forwarding tables of Router 1 / Router 2. The link set corresponds to a specific physical port and its maximum speed limit.

[0049] b. Cache Optimization: To achieve better routing optimization results, network caching must also be optimized. Therefore, the caching optimization problem is designed as a two-step decision process: identifying key nodes and selecting the service types that need to be cached for each key node. The caching scheme can be represented as: .

[0050] in, Indicates at time A defined set of key nodes This represents each key node. This indicates the number of critical nodes, i.e., the number of routers selected in the solution that support caching services. It's worth noting that a larger number... While increasing the upper limit of cache performance, it also brings higher computational overhead and memory usage.

[0051] Key nodes correspond to routers / servers with local storage capabilities: router-type caches use onboard SSDs or memory as content caches; server-type caches are deployed in the server's storage array. Cache hits are returned directly by the device's hardware forwarding plane or accelerated via DPDK.

[0052] c. A model was developed for the joint optimization problem of routing and caching, and optimization was performed. and The joint solution is used to improve QoS performance. Its optimization objective function is:

[0053] in, For the scheduling end time, the utility function This represents a weighted sum of QoS performance metrics, including average end-to-end latency. Average packet loss rate Load balancing coefficient and average route hops .

[0054] Latency and packet loss are measured by the hardware queue length, Egress / Ingress counters, and timestamp registers of each device port; link capacity is determined by the interface line speed and physical layer coding efficiency.

[0055] To further introduce continuous variables With discrete variables To address the nonconvexity of the data, a joint routing and caching algorithm based on deep reinforcement learning is proposed. This algorithm utilizes a single deep reinforcement learning model to simultaneously optimize routing and caching strategies, thereby meeting the needs of fast content distribution services.

[0056] Step 102, Design of a routing optimization algorithm based on deep reinforcement learning.

[0057] Specifically, to improve the speed of content distribution in fusion networks, a routing optimization algorithm based on deep reinforcement learning is proposed. This algorithm includes an offline training phase and an online execution phase.

[0058] a. Offline Training Stage: The Actor Network and Critic Network jointly participate in policy evaluation and optimization; offline training runs on the server and is accelerated by a multi-core GPU; training data is collected through the management port or control VLAN without occupying the service plane.

[0059] b. Execution Phase: After policy convergence, the online execution phase begins, where only the participant network performs inference and publishes routing and caching schemes. Online inference can be performed on the server or edge devices with inference acceleration; policy distribution updates the hardware entries and queue schedulers of Router 1 / Router 2 via NETCONF or OpenFlow.

[0060] Figure 5 The flowchart of the routing optimization algorithm based on deep reinforcement learning includes: Step 201, offline training.

[0061] Specifically, at the start of offline training, the participants' network parameters With commenter network parameters Initialized randomly. The learning records of the strategy are stored using the experience replay pool. The training process consists of... Wheels, each wheel contains One training step. In each round: a. The agent first interacts with the environment to obtain records of each training step and stores them in the experience replay pool; the environmental state is obtained from hardware statistics of switch / router ports and link layer alarm sampling.

[0062] b. The sampling length from the experience replay pool is... The trajectory, in which This trajectory is used to calculate the loss function for participants and commenters; c. To improve sampling efficiency, trajectory sampling and network updates are repeated in each round. Second-rate.

[0063] d. The experience replay pool is cleared after each training round. In the final stage of offline training, the participants' network parameters are retained. .

[0064] Step 202, execute online.

[0065] Specifically, since the participant network and the commentator network are decoupled, only the participant network participates in joint routing and caching decisions during the online execution phase. In real-world scenarios, due to router limitations or unexpected failures, the agent's caching scheme may not be fully deployable on all routers that support caching services. Therefore, the proposed algorithm is compatible. Specifically: a. When the agent infers a joint action If the caching solution cannot be fully deployed, it can be obtained from... Decompose emergency actions Its decomposition form is:

[0066] in, This indicates a key node capable of successfully receiving policies issued by the intelligent agent. Typical scenarios where full deployment fails include degradation triggered by hardware alarms such as insufficient storage slots on the router, NVMe health threshold alarms, port UP / DOWN, and abnormal optical module power.

[0067] b. Other routers A heuristic caching algorithm is employed. In actual deployment, agent actions... Reconfigurable Combined with traditional caching algorithms. During degradation, hardware QoS queues and priority buffers work together with PFC / ECN to ensure critical traffic pass-through.

[0068] c. At the start of online execution, the data learned during the offline training phase... Will be deployed on the participant network : (1) The agent first observes the state of the environment. And reason about actions ; (2) The controller knowledge layer monitors the operation of each cache router: if the caching policy is fully executed, the action is executed directly; if it is not fully executed, it is decomposed into emergency actions and traditional caching policies.

[0069] (3) To perform this action and drive the environment to evolve to .

[0070] The actions performed include: updating the ECMP weights of Router 1 / Router 2, modifying ACLs and forwarding table entries, and adjusting port rates and queue scheduling. The above updates take effect immediately at the ASIC layer and are confirmed asynchronously at the control plane.

[0071] Step 203, Algorithm State Design.

[0072] a. State space It should include necessary information about the network environment. Define the time. status for:

[0073] Wherein, time is represented. link The load, Indicates link The capacity. This is the hardware reading of the port's real-time queue usage, representing the result of the negotiation between the port line speed and the physical layer. b. The first part of the state is the link utilization collected by the routing subproblem; the second part is all data flow characteristics collected by the caching subproblem.

[0074] c. At time Key data streams are identified and fixed through deep reinforcement learning. One key data stream, among which The total number of data streams is always less than the total number of data streams at any given time step. Therefore, its state dimension is... Key data streams are identified on the device through quintuple matching or hardware sampling.

[0075] Step 204, Algorithm Action Design.

[0076] a. For joint routing optimization of multiple data streams, define the time interval. The routing-caching optimization actions are as follows:

[0077] b. The actions of the intelligent agent consist of three parts: setting the link weights, identifying the set of key cache nodes, and selecting the type of cached data stream.

[0078] c. At each timestamp, the agent needs to identify There are several key nodes, therefore the action dimension is fixed as follows: .

[0079] The link weight is set according to the ECMP weight and queue scheduling parameters; the key nodes are devices with sufficient memory and idle PCIe channels; the data flow type is assigned to different QoS queues through a hardware classifier.

[0080] Step 205, design the algorithm reward function.

[0081] reward function The evaluation of the effectiveness of the agent's policy learning must be consistent with the optimization objective of the objective function in step 101: a. Incorporate QoS metrics (end-to-end latency) Packet loss rate Load balancing coefficient With route hop count This is converted into the corresponding reward value. The specific calculation formula is as follows:

[0082]

[0083]

[0084]

[0085] in, Indicates time Maximum end-to-end latency of all data streams This represents the maximum number of hops for all possible paths in the network, therefore A higher reward value corresponds to better QoS performance. Latency is estimated by hardware timestamp / queue dwell time; packet loss rate is derived from port and queue drop counters. b. Total reward value Defined as a linear weighted sum of the individual rewards:

[0086] in, satisfy:

[0087] therefore , The larger the value, the better the QoS performance in terms of end-to-end latency, packet loss rate, load balancing, and routing hop count, thereby achieving the optimization objective of the objective function in step 101.

[0088] As can be seen, the routing optimization problem modeled in this disclosure, in a cross-domain heterogeneous network environment, transforms multiple QoS constraints into optimization objectives and constructs a state space, action space, and transition probabilities, clarifying the decision inputs and outputs of the agent in a dynamic network environment. The routing optimization based on deep reinforcement learning in this disclosure uses a deep neural network to extract features from the state space and utilizes a multi-objective reward function to jointly evaluate latency optimization, link load balancing, and energy consumption control, thereby achieving dynamic updates of the routing strategy and optimal path selection.

[0089] In step S330, the transmission association information of the data to be transmitted is obtained. The transmission association information is used to indicate the data information, data link information and data scheduling information of the data to be transmitted.

[0090] Specifically, by extracting multi-dimensional transmission correlation information (including data characteristic attributes, transmission path configuration, and scheduling strategy parameters) from the data to be transmitted, a decision-making basis for intelligent network transmission is provided. Specifically, data information includes metadata such as data type, volume, and priority, used for differentiated service strategy formulation; data link information integrates status parameters such as path topology, real-time load, and available bandwidth to support dynamic path selection; and data scheduling information combines timing constraints and resource allocation rules to optimize transmission timing. Through global perception and collaborative analysis of this information, the system can achieve resource-aware intelligent routing (such as avoiding congested links), differentiated quality of service assurance (prioritizing high-priority data scheduling), and maximize transmission efficiency (matching data flows based on link characteristics), effectively improving network throughput and reducing end-to-end latency, making it particularly suitable for efficient and reliable transmission scenarios in complex and dynamic network environments.

[0091] In step S340, based on the transmission association information, data transmission information that minimizes transmission time while satisfying preset constraints is determined.

[0092] This approach involves deeply integrating the associated information of the data to be transmitted (covering data characteristics, link status, and scheduling rules) to construct a multi-constraint optimization model to solve for the optimal decision that minimizes transmission time. Under preset conditions such as bandwidth capacity, quality of service assurance, and cache resource limitations, dynamic programming or heuristic algorithms are used to jointly optimize data fragmentation strategies, transmission path selection, and cache scheduling timing. This ensures that high-priority data is scheduled first, link load is balanced, and cache resources are reused efficiently. This can shorten end-to-end transmission latency, maximize resource utilization while ensuring transmission reliability, and is particularly suitable for the efficient carrying of latency-sensitive services (such as industrial real-time control and high-definition video live streaming), effectively balancing the contradiction between transmission efficiency and system constraints.

[0093] In some embodiments, in response to the transmission association information, data scheduling information that minimizes transmission duration when satisfying preset constraints is determined, including: Based on the trained data transmission model, feature extraction is performed on the transmission association information to obtain transmission association features; and based on the transmission association features, the data scheduling information is determined. The process of training the initial model based on training samples to obtain the trained data transmission model includes: The sample constraints of the training samples are determined, including priority constraints and maximum transmission time constraints; the training samples include historical data transmission information, historical transmission duration, and corresponding transmission reward functions. Feature extraction is performed on the historical data transmission information to obtain historical data transmission features; Based on the historical data transmission characteristics, the predicted transmission duration that satisfies the sample constraints is determined; The model parameters of the initial model are updated based on the transmission reward function to minimize the loss function between the predicted transmission duration and the historical transmission duration, thereby obtaining the trained data transmission model.

[0094] This approach achieves efficient data scheduling optimization under constraints by constructing an intelligent data transmission model. The model is trained using historical transmission data, and during the feature extraction stage, sample information such as priority constraints and maximum transmission latency limits is integrated. The model parameters are dynamically adjusted to minimize prediction loss through feedback from reward functions comparing predicted and actual transmission times, enabling the model to grasp latency optimization patterns under multiple constraints. In practical applications, the model extracts key features based on real-time transmission correlation information and quickly generates optimal scheduling strategies that meet preset priority and time limit requirements. This combination of reinforcement learning and constraint optimization ensures both the transmission priority and timeliness of critical data, while also adapting to dynamic network changes through a data-driven model, significantly reducing end-to-end transmission latency and effectively improving resource scheduling efficiency and service quality assurance capabilities in complex network scenarios.

[0095] In some embodiments, the data information includes the dynamic priority and remaining data volume of the data to be transmitted; the data link information includes the resource utilization rate of the data link, the current link congestion level and available resources; and the data scheduling information includes data allocation information between data links and the corresponding data transmission time. And / or, the transmission reward function includes: ;in, For the first Data The transmission reward function, For load balancing capabilities, For the first Data Dynamic priority.

[0096] This scheme integrates the dynamic priority and remaining capacity of data to be transmitted with the real-time resource status of links (utilization, congestion level, and available resources) to construct a scheduling mechanism based on load balancing and priority co-optimization. A transmission reward function designed using dynamic priority (MTT) and a load balancing factor (β) ensures that high-priority data is transmitted first during data allocation, while also achieving link load balancing by penalizing excessive resource concentration. The data allocation strategy is dynamically adjusted based on available link resources and congestion levels, and the transmission time of each link is accurately calculated. This scheme significantly improves scheduling adaptability in complex network environments, effectively reducing end-to-end latency and avoiding local link overload while ensuring the timeliness of critical data. Through dual-objective optimization of efficient resource utilization and load balancing, it achieves a comprehensive improvement in overall transmission efficiency and stability.

[0097] Specifically, a data transmission model based on deep reinforcement learning can be designed. This model utilizes deep reinforcement learning algorithms for transmission scheduling, enabling it to dynamically adjust transmission strategies according to changes in the transmission environment and service requirements, thereby achieving real-time, efficient, and reliable data transmission. The core idea of ​​the algorithm is to introduce maximum entropy reinforcement learning to ensure its efficiency, stability, and robustness. Transmission scheduling directly affects the hardware queuing, rate shaping, and congestion control of the switch / router's outgoing ports, and combines this with multi-queue and DMA ring buffer optimization on the server-side NIC to achieve end-to-end performance guarantees.

[0098] Discrete action space typically refers to an action space with a finite number of specific actions. Since this data transmission model primarily addresses the transmission scheduling problem on a finite number of data links, this invention appropriately improves the policy function to output a probability distribution of actions within the discrete action space, thereby ensuring the algorithm's applicability in this space. Figure 6 This is a flowchart of a network data transmission algorithm based on deep reinforcement learning.

[0099] Step 301, State space design of network data transmission algorithm.

[0100] The state space of a network data transmission algorithm refers to the set of all possible states that an agent can exist in its environment. A state refers to the specific situation in which the agent is situated in the environment. These states affect the agent's next action, and the agent's actions, in turn, change its current state.

[0101] Specifically, this network data transmission algorithm will use the state space... The design consists of three parts: a data information set, a data link information set, and a scheduling information set. The calculation formula is as follows:

[0102] in, Represents a set of data information. Represents the data link information set. Represents the scheduling information set: a. Data information set Includes dynamic priority of transmitted data and remaining data volume Dynamic priority This reflects the importance of the data, business needs, and latency during transmission; remaining data volume. This indicates the size of the portion of each data file that has not yet been transferred. Before the transfer begins... Equal to the total size of the data file; when the transfer is complete, The data resides on the server's local NVMe or in memory buffer. The NIC sends data in fragments via DMA, and the fragment size is affected by the MTU and the hardware's segmentation capabilities.

[0103] b. Link Information Set Includes resource utilization of each data link The system displays the current link congestion level and available resources. By calculating link resource utilization in real time, corresponding scheduling and optimization can be performed. Resource utilization is derived from port rate counters and queue occupancy levels; congestion level is referenced to ECN tagging rate and queue latency; available resources are determined by port negotiation rate and reserved bandwidth (hardware shaper).

[0104] c. Scheduling information set Includes data allocation between links and data transmission time The scheduling status reflects which link data has been allocated to for transmission and which data has not yet been scheduled. Allocation corresponds to a specific output port and a hardware queue mapping, and transmission time is measured by hardware timestamps and port queuing delays.

[0105] Step 302, Action space design of network data transmission algorithm.

[0106] Discrete action space typically refers to a finite set of discrete actions. An action is an operation taken by an agent in a specific state to maximize its cumulative reward. Through interaction with the environment and continuous learning, the agent can continuously update and optimize its behavioral strategies, thereby achieving optimal decision-making.

[0107] a. In this data transmission model, the action space Designed to represent all data links, its calculation formula is as follows:

[0108] in, Indicates a data link. This indicates the number of data links. Each data link corresponds one-to-one with a device's physical port or aggregated link, and each link has its own independent hardware queue and shaping parameters.

[0109] b. To achieve optimal data transmission, multiple factors need to be considered when setting up data links. Too many links may increase system resource consumption and algorithm overhead, while too few links may lead to excessively long data streams, unbalanced loads, and excessively high transmission latency.

[0110] c. When configuring the link, factors such as data priority, bandwidth, and latency need to be considered comprehensively to achieve optimal data transmission performance. Priority is bound to the hardware queue via DSCP / 802.1p tags; bandwidth is limited by port line speed and shaper rate; latency is affected by physical distance (fiber length), encoding / decoding latency, and queue dwell time.

[0111] Step 303, Design of the reward function for the network data transmission algorithm.

[0112] The reward function measures the reward an agent receives for taking an action in a given state. This mechanism enables the agent to learn and optimize its policy through continuous interaction with the environment, thereby achieving optimal data transmission scheduling.

[0113] The optimization objective of this model is to achieve load balancing and maintain the maximum transmission time (MTT) of the dataset while satisfying dynamic data priority. Therefore, the reward function is designed as follows:

[0114] in, Indicates load balancing capability. Indicates the first Data sets, Indicates the first Dynamic priority of each dataset.

[0115] As shown in the formula above, the reward value increases when data has a higher priority and a shorter transmission time. This design guides agents to prioritize the scheduling of important data, thereby improving the performance and efficiency of the overall data transmission model.

[0116] Load balancing at the physical layer is manifested as parallel outflow from multiple ports, achieved through hardware hashing or weight allocation. When a port failure is detected (such as optical power or LOS alarm), the hardware quickly converges and triggers policy recalculation to ensure that the reward target is not affected by a single point of failure.

[0117] In step S350, a routing action is performed based on the joint policy information, and a data scheduling action is performed based on the data transmission information to transmit the data to be transmitted.

[0118] This system achieves intelligent data transmission control by collaboratively utilizing joint policy information and data transmission information: It dynamically selects the optimal routing path based on link weights in the joint policy to avoid congested areas, while pre-storing high-frequency data stream types at key cache nodes to reduce redundant transmission distances; combining dynamic priorities and resource allocation rules in the data transmission information, it precisely schedules transmission timing according to data urgency and available link resources, ensuring the timeliness of high-priority data and balancing the overall network load. This improves transmission efficiency (shortening end-to-end latency), enhances service reliability (reducing packet loss risk through load balancing), and maximizes resource utilization (coordinated optimization of link capacity and cache space). It exhibits adaptive scheduling capabilities, especially in dynamically changing network environments, comprehensively ensuring the real-time performance, stability, and user experience of data transmission.

[0119] In some embodiments, performing routing actions based on the federated policy information includes: In response to the determination that the critical cache node cannot fully deploy the federated policy information, a heuristic caching algorithm is used to perform caching operations; In response to the detection that the critical cache node can fully deploy the federated policy information, a cache operation is performed based on the federated policy information; Specifically, in response to detecting a hardware anomaly in the critical cache node, it is determined that the critical cache node cannot fully deploy the federated policy information; the hardware anomaly includes at least one of the following: insufficient storage slots on the router, non-volatile memory host controller interface specification health threshold alarm, abnormal port startup / shutdown, and abnormal optical module power.

[0120] By dynamically assessing the deployment capabilities and hardware health of key cache nodes, the system achieves elastic execution of routing and caching policies. When hardware limitations such as insufficient router slots, NVMe health alarms, port anomalies, or optical module power anomalies are detected, the system automatically switches to a heuristic caching algorithm, dynamically adjusting data cache locations and allocation strategies based on resource availability to ensure the executability of some policies. If the node hardware is normal and resources are sufficient, the cache is deployed strictly according to the joint policy, prioritizing efficient storage and fast retrieval of high-value data. This enhances the robustness of the network system, maintaining basic service continuity even in the event of hardware failure or resource constraints. Furthermore, through dual-mode adaptation of policies and algorithms, cache resource utilization is optimized and transmission latency is reduced. Combined with proactive monitoring of hardware anomalies, fault prevention and rapid response are achieved, comprehensively improving data transmission reliability in complex scenarios.

[0121] In some embodiments, performing routing actions based on the federated policy information includes: It takes effect immediately at the application-specific integrated circuit (ASIC) layer and is asynchronously confirmed at the control plane for at least one of the following: updating the router's link weight, modifying access control lists and forwarding table entries, and adjusting port rates and queue scheduling policies. And / or, obtain transmission association information of the data to be transmitted, including: The dynamic priority of the data to be transmitted is determined based on at least one of the importance of the data to be transmitted, business requirements, and the allowable waiting time during the transmission process. The untransmitted portion of the data to be transmitted is detected to obtain the remaining data volume; The resource utilization rate is determined based on the port rate counter and the queue occupancy level. The degree of link congestion is determined based on the explicit congestion notification marking rate and queue delay; The available resources are determined based on the port negotiation rate and the reserved bandwidth. The data allocation information is determined based on the mapping relationship between the output port and the hardware queue; The data transmission time is determined based on hardware timestamps and port queuing delays.

[0122] This system achieves efficient routing and data scheduling through hardware and software co-optimization: Joint policies (such as link weight updates and forwarding table modifications) are executed in real-time at the ASIC layer, ensuring low-latency hardware-level response, while asynchronous acknowledgment in the control plane guarantees policy consistency. By combining dynamic priorities (based on service requirements and waiting time), real-time resource status (port rate, queue occupancy, congestion notification), and hardware characteristics (queue mapping, timestamps), data allocation and transmission time are accurately calculated. This approach leverages hardware acceleration to improve policy execution efficiency while intelligent decision-making at the software layer optimizes resource allocation, significantly reducing end-to-end latency and enhancing the network's adaptability to sudden traffic surges and dynamic loads. Furthermore, priority-aware scheduling ensures the quality of service for critical businesses, comprehensively improving transmission reliability and resource utilization in complex network environments.

[0123] As can be seen, the data transmission optimization based on deep reinforcement learning disclosed herein dynamically adjusts the data packet transmission rate, scheduling order, and congestion control mechanism by real-time monitoring of link status and service flow characteristics, thereby achieving differentiated protection for different QoS services and improving network throughput and data transmission stability.

[0124] Therefore, this disclosure provides an efficient and intelligent routing optimization and data transmission method suitable for triple-play environments, which can solve the problem that traditional solutions struggle to dynamically meet multiple Quality of Service (QoS) requirements and achieve effective load balancing. Addressing the poor adaptability and singular optimization objectives of existing routing algorithms in cross-domain, multi-protocol converged networks, this paper proposes a routing optimization and network data transmission method based on deep reinforcement learning. By constructing an interactive learning framework between the agent and the environment, it achieves collaborative dynamic optimization of routing strategies and data transmission. A single deep reinforcement learning model is used to jointly optimize multiple QoS indicators, and a multi-objective reward function coordination mechanism is introduced, significantly enhancing the overall load balancing capability of the system while improving content distribution efficiency.

[0125] It should be noted that the method of this embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this embodiment, and the multiple devices will work together to generate video to complete the method described.

[0126] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0127] Based on the same technical concept, corresponding to any of the methods in the above embodiments, this disclosure also provides a network data transmission apparatus for routing optimization based on deep reinforcement learning, see [link to relevant documentation]. Figure 7 The network data transmission device for routing optimization based on deep reinforcement learning includes: An environment acquisition module is used to acquire environmental status information of the network environment, wherein the environmental status information is used to indicate link utilization information and data flow information of the network environment; The joint optimization module is used to jointly optimize routing and caching strategies based on the environmental state information to obtain joint strategy information about routing and caching; the joint strategy information includes link weights, key cache nodes, and cache data stream types. The transmission information acquisition module is used to acquire transmission association information of the data to be transmitted, wherein the transmission association information is used to indicate the data information, data link information and data scheduling information of the data to be transmitted. The transmission optimization module is used to determine, based on the transmission association information, the data transmission information that minimizes the transmission time while satisfying preset constraints; The action execution module is used to execute routing actions based on the joint policy information and data scheduling actions based on the data transmission information to transmit the data to be transmitted.

[0128] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0129] The apparatus of the above embodiments is used to implement the network data transmission method based on deep reinforcement learning for routing optimization in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0130] Based on the same technical concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the network data transmission method for routing optimization based on deep reinforcement learning as described in any of the above embodiments.

[0131] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0132] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the network data transmission method based on deep reinforcement learning for routing optimization as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0133] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.

[0134] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0135] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0136] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A network data transmission method based on deep reinforcement learning for routing optimization, comprising: Obtain environmental status information of the network environment, wherein the environmental status information is used to indicate link utilization information and data flow information of the network environment; Based on the environmental state information, a joint optimization of routing and caching strategies is performed to obtain joint strategy information for routing and caching; the joint strategy information includes link weights, key cache nodes, and cache data stream types; Obtain transmission association information of the data to be transmitted, wherein the transmission association information is used to indicate the data information, data link information and data scheduling information of the data to be transmitted; Based on the transmission association information, determine the data transmission information that minimizes the transmission time while satisfying preset constraints; Based on the joint policy information, routing actions are performed, and based on the data transmission information, data scheduling actions are performed to transmit the data to be transmitted.

2. The method according to claim 1, wherein, The environmental status information includes: ; in, For at any time Environmental status information, For at any time Link The load, For at any time Link capacity, for The weight, for The weight, For at any time The hardware readings for the real-time queue usage of node m's port. For at any time The port line speed of node m is negotiated with the physical layer. Based on the aforementioned environmental state information, a joint optimization of routing and caching strategies is performed to obtain joint strategy information regarding routing and caching, including: Based on the trained joint optimization network, feature extraction is performed on the environmental state information to obtain environmental state features; and based on the environmental state features, the link weights of links and key cache nodes in the network environment are determined. The corresponding cache data stream type is determined based on the hardware type of the key cache node; Based on the link weight, the key cache node, and the cache data stream type, the joint policy information at time u is determined; wherein, the joint policy information includes: ;in, For the joint policy information at time u, For the link at time u The link weights, where F is the set of links. For the critical cache node at time u, At time u, in the cache node The set of nodes at time u For the key cache node at time u The cached data stream type.

3. The method according to claim 2, further comprising: The participant network and the commenter network are trained based on historical policy learning records, and the trained participant network is determined as the joint optimization network; wherein, Historical learning trajectory information of a preset length is sampled from the historical strategy learning records, and features are extracted from the historical learning trajectory information to obtain historical learning trajectory features; The joint reward function and the loss function between the participant network and the commentator network are determined based on the historical learning trajectory features. The network parameters of the initial joint network are updated based on the joint reward function to minimize the loss function, thereby obtaining the trained participant network. The joint reward function includes: Among them, the end-to-end latency reward value , For at any time Maximum end-to-end latency of all data streams For at any time End-to-end latency; packet loss rate reward value , For at any time Packet loss rate; Load balancing reward value , For at any time Load balancing coefficient; hop count bonus , For at any time The number of hops in the path. This represents the maximum number of hops across all possible paths in the network environment. The weights for end-to-end latency reward values, The weight of the packet loss rate reward value, The weights for load balancing reward values. The weight of the route hop count reward value; The link weights and nodes that maximize the joint reward function are determined as the link weights and the critical cache nodes.

4. The method according to claim 1, wherein, Executing routing actions based on the aforementioned joint policy information includes: In response to the determination that the critical cache node cannot fully deploy the federated policy information, a heuristic caching algorithm is used to perform caching operations; In response to the detection that the critical cache node can fully deploy the federated policy information, a cache operation is performed based on the federated policy information; Specifically, in response to detecting a hardware anomaly in the critical cache node, it is determined that the critical cache node cannot fully deploy the federated policy information; the hardware anomaly includes at least one of the following: insufficient storage slots on the router, non-volatile memory host controller interface specification health threshold alarm, abnormal port startup / shutdown, and abnormal optical module power.

5. The method according to claim 1, wherein, Based on the transmission association information, data scheduling information that minimizes transmission time while satisfying preset constraints is determined, including: Based on the trained data transmission model, feature extraction is performed on the transmission association information to obtain transmission association features; and based on the transmission association features, the data scheduling information is determined. The process of training the initial model based on training samples to obtain the trained data transmission model includes: The sample constraints of the training samples are determined, including priority constraints and maximum transmission time constraints; the training samples include historical data transmission information, historical transmission duration, and corresponding transmission reward functions. Feature extraction is performed on the historical data transmission information to obtain historical data transmission features; Based on the historical data transmission characteristics, the predicted transmission duration that satisfies the sample constraints is determined; The model parameters of the initial model are updated based on the transmission reward function to minimize the loss function between the predicted transmission duration and the historical transmission duration, thereby obtaining the trained data transmission model.

6. The method according to claim 5, wherein, The data information includes the dynamic priority and remaining data volume of the data to be transmitted; the data link information includes the resource utilization rate of the data link, the current link congestion level and available resources; and the data scheduling information includes the data allocation information between data links and the corresponding data transmission time. And / or, the transmission reward function includes: ;in, For the first Data The transmission reward function, For load balancing capabilities, For the first Data Dynamic priority.

7. The method according to claim 6, wherein, Executing routing actions based on the aforementioned joint policy information includes: It takes effect immediately at the application-specific integrated circuit (ASIC) layer and is asynchronously confirmed at the control plane for at least one of the following: updating the router's link weight, modifying access control lists and forwarding table entries, and adjusting port rates and queue scheduling policies. And / or, obtain transmission association information of the data to be transmitted, including: The dynamic priority of the data to be transmitted is determined based on at least one of the importance of the data to be transmitted, business requirements, and the allowable waiting time during the transmission process. The untransmitted portion of the data to be transmitted is detected to obtain the remaining data volume; The resource utilization rate is determined based on the port rate counter and the queue occupancy level. The degree of link congestion is determined based on the explicit congestion notification marking rate and queue delay; The available resources are determined based on the port negotiation rate and the reserved bandwidth. The data allocation information is determined based on the mapping relationship between the output port and the hardware queue; The data transmission time is determined based on hardware timestamps and port queuing delays.

8. A network data transmission device for routing optimization based on deep reinforcement learning, comprising: An environment acquisition module is used to acquire environmental status information of the network environment, wherein the environmental status information is used to indicate link utilization information and data flow information of the network environment; The joint optimization module is used to jointly optimize routing and caching strategies based on the environmental state information to obtain joint strategy information about routing and caching; the joint strategy information includes link weights, key cache nodes, and cache data stream types. The transmission information acquisition module is used to acquire transmission association information of the data to be transmitted, wherein the transmission association information is used to indicate the data information, data link information and data scheduling information of the data to be transmitted. The transmission optimization module is used to determine, based on the transmission association information, the data transmission information that minimizes the transmission time while satisfying preset constraints; The action execution module is used to execute routing actions based on the joint policy information and data scheduling actions based on the data transmission information to transmit the data to be transmitted.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as claimed in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method of any one of claims 1 to 7.