Intelligent computing fusion network system and computing network resource adaptation and reliable transmission scheduling method

Through intelligent computing fusion network architecture and cross-domain orchestration and scheduling methods, the problems of privacy and communication overhead in AIGC model training are solved, and data privacy protection and efficient and reliable model training and transmission are achieved.

CN119892834BActive Publication Date: 2025-10-21BEIJING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510052587.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-10-21
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

In industrial production, the training of AIGC models requires massive data samples, which leads to privacy leakage and excessive communication overhead, making it difficult to achieve real-time and accurate production control and decision support.

Method used

It adopts an intelligent computing fusion network architecture, including the terminal device layer, edge service layer and cloud control layer, performs distributed computing through the federated learning paradigm, combines the particle swarm algorithm and deep reinforcement learning algorithm for cross-domain orchestration and traffic scheduling, and uses time-sensitive network and deterministic network technology for reliable transmission scheduling.

Benefits of technology

It achieves data privacy protection and communication overhead reduction of the AIGC model, improves model training efficiency and accuracy, and ensures reliable transmission of parameter streams and efficient end-to-end collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119892834B_ABST
    Figure CN119892834B_ABST
Patent Text Reader

Abstract

The application discloses a kind of wisdom calculation fusion network architecture and algorithm network resource adaptation and reliable transmission scheduling method, it is related to industrial communication network field, this wisdom calculation fusion network architecture includes terminal equipment layer, edge service layer and cloud control layer, by the training process of AIGC model being migrated to each ICPS domain computing device from cloud server by federated learning paradigm, distributed computing device trains local model according to local data set, then by exchanging model parameters with cloud server to realize the iterative optimization of global model, this kind of mode guarantees data privacy, and also reduces communication overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of industrial communication networks, and in particular to an intelligent computing fusion network architecture and a computing network resource adaptation and reliable transmission scheduling method. Background Art

[0002] With the booming development of artificial intelligence-generated content (AIGC) technology, AIGC models have shown broad application prospects in the industrial field. For example, AIGC models can learn from large amounts of equipment operation data and generate equipment health status samples, which can be used to train and test fault detection algorithms to improve detection accuracy. Therefore, industrial cyber-physical systems (ICPS) are expected to achieve more precise production control and decision support by building AIGC services. However, the industrial production process requires real-time and precise control and scheduling of on-site equipment, which places more stringent requirements on the quality and response time of AIGC services. Building AIGC services also faces a series of challenges. First, AIGC model training requires massive data samples, which may lead to privacy leaks and huge communication overhead during the process. Summary of the Invention

[0003] The purpose of this application is to provide an intelligent computing fusion network architecture and a computing network resource adaptation and reliable transmission scheduling method, which not only ensures data privacy during AIGC model training, but also reduces communication overhead.

[0004] To achieve the above objectives, this application provides the following solutions:

[0005] In a first aspect, the present application provides an intelligent computing fusion network architecture, including:

[0006] The terminal device layer includes a plurality of industrial cyber-physical system domains distributed in different regions or factories, each of which includes a computing device, each of which is used to train a local model based on a local sample data set and target parameters to obtain local model training parameters, wherein the target parameters are the AIGC model parameters that need to be updated and are allocated by the cloud control layer to each of the industrial cyber-physical system domains;

[0007] An edge service layer, connected to the terminal device layer, for performing edge aggregation processing on the local model training parameters to obtain edge model parameters;

[0008] The cloud control layer is connected to the edge service layer through the core network layer, and is used to perform global aggregation processing on the edge model parameters to obtain global model parameters.

[0009] In a second aspect, the present application provides a computing network resource adaptation and reliable transmission scheduling method, which applies the intelligent computing fusion network architecture described in the first aspect. The computing network resource adaptation and reliable transmission scheduling method includes:

[0010] The cloud control layer uses a particle swarm algorithm to generate an optimal cross-domain orchestration solution based on the computing resources of each industrial cyber-physical system domain and the training requirements of the AIGC model. The optimal cross-domain orchestration solution includes the industrial cyber-physical system domain selection result and the computing resource allocation ratio of the corresponding industrial cyber-physical system domain.

[0011] The terminal device layer determines the target industrial cyber-physical system domain according to the optimal cross-domain orchestration solution, and the target industrial cyber-physical system domain trains a local model according to the local sample data set and the allocation parameters to obtain local model training parameters;

[0012] The edge service layer performs edge aggregation processing on the local model training parameters to obtain edge model parameters;

[0013] The cloud control layer performs global aggregation processing on the edge model parameters to obtain global model parameters.

[0014] According to the specific embodiments provided in this application, this application has the following technical effects:

[0015] The present application provides an intelligent computing fusion network architecture and a method for computing network resource adaptation and reliable transmission scheduling. The intelligent computing fusion network architecture includes a terminal device layer, an edge service layer, and a cloud control layer. Through the federated learning paradigm, the training process of the AIGC model is migrated from the cloud server to the computing devices of each ICPS domain. The distributed computing devices train local models based on local data sets, and then achieve iterative optimization of the global model by exchanging model parameters with the cloud server. This method not only ensures data privacy but also reduces communication overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 A schematic diagram of an intelligent computing fusion network architecture provided in Example 1 of the present application;

[0018] Figure 2 This is a flowchart of the cross-domain computing orchestration algorithm based on particle swarm optimization in Example 1 of the present application;

[0019] Figure 3 This is a flowchart of a deterministic transmission scheduling algorithm based on proximal policy optimization in Example 1 of the present application;

[0020] Figure 4 This is a flowchart of the multi-queue circular queuing and forwarding mechanism in Example 1 of this application. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0023] Example 1

[0024] This embodiment provides an intelligent computing fusion network architecture to support the efficient construction of AIGC services in ICPS, such as Figure 1 As shown in the figure, the intelligent computing fusion network architecture includes four layers: terminal device layer, edge service layer, core network layer, and cloud control layer. The coordination and cooperation between the layers realize cross-domain computing collaboration and end-to-end reliable transmission scheduling of parameter streams.

[0025] The terminal device layer is responsible for managing local data sample sets and local models, training and updating local models according to the parameters issued by the edge service layer, realizing distributed training of the AIGC model and uploading local model parameters to the edge service layer. The terminal device layer includes several industrial information-physical system domains distributed in different regions or factories. Each of the industrial information-physical system domains includes a computing device. Each of the computing devices is used to train the local model according to the local sample data set and target parameters to obtain local model training parameters, wherein the target parameters are the AIGC model parameters that need to be updated and assigned by the cloud control layer to each of the industrial information-physical system domains.

[0026] During the training phase of the AIGC model, the terminal device layer adopts the distributed computing paradigm of federated learning, migrating the training process from a centralized computing center to various ICPS domains. These domains are located in different regions or factories, each containing industrial equipment and computing devices used for model training and managing the sample datasets and local models they hold. Multiple ICPS domains collaborate to complete the training task of the same large industrial model. Specifically, the ICPS domain trains and updates the local model based on the sample dataset and parameters (i.e., target parameters) issued by the edge service layer. The updated local model parameters are then uploaded to the edge service layer for edge aggregation, enabling distributed collaborative training of the AIGC model and improving the efficiency and accuracy of model training.

[0027] The terminal device layer is connected to the edge service layer through a switch. Specifically, the terminal device layer is associated with the edge service layer through a time-sensitive networking (TSN) switch. The TSN switch implements a cyclic queuing and forwarding (CQF) mechanism. It uses two queues (i.e., switch port queues) to alternately receive and send parameter streams, and switches the queue transmission and reception states based on a pre-configured cyclic time slot based on a switch gating list, thereby providing deterministic transmission guarantees for parameter streams and achieving reliable transmission scheduling. The switch gating list represents the open and closed states of the switch queue sending gate and the switch queue receiving gate in each preset cyclic time slot. The switch queue sending gate is open when the switch port queue sends the parameter stream, and the switch queue receiving gate is open when the switch port queue receives the parameter stream.

[0028] The edge service layer is used to connect terminal devices to the core network layer, and is responsible for performing edge aggregation of local model parameters to form edge model parameters, and uploading the edge model parameters to the cloud control layer to reduce the traffic transmission burden of the core network layer.

[0029] The edge service layer includes several device groups, each of which includes access devices and edge servers; the access devices are composed of access routers and wireless access points (such as base stations) located at the edge of the network, and are used to connect terminal devices to the core network layer, wherein the terminal devices include computing devices, and the access devices undertake ICPS domains and wide-area deterministic networks (WADN); each access device covers a certain number of ICPS domains, and the corresponding edge server performs edge aggregation on the local model parameters of the ICPS domains within its coverage area to form edge model parameters, which are uploaded to the cloud server through WADN for global aggregation.

[0030] The core network layer deploys a multi-queue round-robin forwarding mechanism in forwarding devices such as routers to provide bounded latency, jitter, and packet loss rate for parameter stream transmission, thereby achieving reliable transmission scheduling in the wide area network.

[0031] The core network layer mainly includes forwarding devices such as WADN routers. Each forwarding device is responsible for the parameter flow transmission between the cloud control layer and the edge service layer based on a multi-queue circular queuing forwarding mechanism. The multi-queue circular queuing forwarding mechanism refers to controlling the working status of several forwarding device port queues in each transmission cycle through a forwarding device gating list. The forwarding device port queue refers to several queues set at the port of each forwarding device. The working status includes receiving parameter flow and sending parameter flow. In each transmission cycle, only one forwarding device port queue has a working status of sending parameter flow, and the working status of the remaining forwarding device port queues is receiving parameter flow. The forwarding device gating list represents the opening and closing status of the forwarding device port queue sending gate and the forwarding device port queue receiving gate in each transmission cycle. The forwarding device port queue sending gate is open when the forwarding device port queue sending gate is open, and the forwarding device port queue receiving gate is open when the forwarding device port queue receiving gate is open when the forwarding device port queue receiving gate is open.

[0032] The time difference between the start times of the same transmission cycle of two adjacent forwarding devices is equal to the link delay between the forwarding devices.

[0033] A multi-queue round-robin forwarding mechanism deployed in forwarding devices effectively addresses the issue of misaligned transmission cycles caused by inter-device link latency. Furthermore, this mechanism enables unified scheduling of parameter streams within the same transmission cycle, even in the presence of clock synchronization deviations. As a result, the core network layer provides bounded latency, jitter, and packet loss for parameter stream transmission, enabling reliable transmission scheduling within the wide-area core network.

[0034] The cloud control layer is connected to the edge service layer through the core network layer, and is used to perform global aggregation processing on the edge model parameters to obtain global model parameters.

[0035] Considering the federated training paradigm of the AIGC model, the data sample size and characteristics of each ICPS domain are different. In addition, the geographically dispersed computing resources are unevenly distributed and relatively independently managed. Inappropriate training equipment selection and computing resource allocation will lead to reduced model training efficiency. In addition, during the training process, model parameters are constantly interacting between distributed devices and cloud servers. By introducing time-sensitive networks and deterministic network technologies, network resources (such as time slots, queues, bandwidth, etc.) are planned and scheduled to achieve deterministic and reliable transmission of parameter streams. However, the transmission of parameter streams generally needs to span multiple heterogeneous networks (such as industrial intranets and wide-area core networks). The network scale is large and the transmission distance is long. A single deterministic transmission mechanism is difficult to guarantee end-to-end reliable transmission scheduling from a global perspective.

[0036] In this regard, the cloud control layer is also used for dynamic adaptation of computing network resources, specifically including using a particle swarm algorithm to generate an optimal cross-domain orchestration plan based on the computing power resources of each industrial information-physical system domain and the training requirements of the AIGC model, wherein the optimal cross-domain orchestration plan includes the industrial information-physical system domain selection result and the computing power resource allocation ratio of the corresponding industrial information-physical system domain; based on the parameter flow information to be scheduled and the network status information, a deep reinforcement learning algorithm is used to generate an optimal traffic scheduling decision, and based on the optimal traffic scheduling decision, the parameter flow transmission path is determined and a gating list of each network communication device on the parameter flow transmission path is configured, wherein the parameter flow information includes an ID number, a flow period, a source address, a destination address, an optional transmission path, a frame size, an end-to-end delay requirement and a jitter requirement, the network status information includes the link bandwidth and the port queue resource capacity of the network communication device, the network communication device includes the forwarding device in the core network, the switch in the terminal device layer and the access device in the edge service layer, and the gating list includes the above-mentioned switch gating list control and the forwarding device gating list.

[0037] The cloud control layer deploys cloud servers and cloud controllers, which are used to achieve global aggregation of AIGC model training parameters and dynamic adaptation of computing network resources respectively.

[0038] The cloud server receives edge model parameters transmitted by the WADN, performs global aggregation, and updates the global model. The updated global model parameters are distributed again via the core network layer and edge service layer to the ICPS domains participating in model training, where they are continuously iterated until the global model reaches the desired accuracy. At the start of model training, the cloud server is responsible for initializing the global model parameters and distributing them to the ICPS domains for iterative training.

[0039] On the one hand, the cloud controller is responsible for coordinating and managing computing resources at the terminal device layer and making cross-domain computing orchestration decisions. This involves selecting the appropriate ICPS domain for training large industrial models and making computing resource allocation decisions to optimize model training latency. On the other hand, it needs to perceive the global network state and make traffic scheduling decisions. This involves allocating time slots (queues) to parameter flows, ensuring deterministic transmission from source to destination, and guaranteeing a high scheduling success rate, thereby achieving reliable transmission scheduling.

[0040] like Figure 2 As shown in Figure 1, the cross-domain computing orchestration algorithm based on particle swarm optimization (i.e., using particle swarm optimization to generate the optimal cross-domain orchestration solution) includes the following steps:

[0041] Step S101: Sense computing resources and training requirements, and generate an initial cross-domain computing orchestration plan. Specifically, the cloud controller dynamically senses the computing resource capacity of each ICPS domain. Sample data volume And the latency requirements for industrial large model training Among them, i u Represents the i-th ICPS domain covered by edge server u; generates the initial cross-domain computing orchestration plan Among them, α u,i (m)∈{0,1} represents ICPS domain i u Whether to participate in training model m, o u,i (m)∈[0,1] represents the proportion of computing resources allocated for training model m.

[0042] Step S102: Initialize the population and define the particle position and velocity. Specifically, generate the initial population according to the initial cross-domain computing arrangement scheme. The population Including several particles p k , where k represents the kth particle in the population; each of the particles p k Contains position p k,t and speed v k,t Two elements, where t represents the tth iteration; each particle position corresponds to an initial cross-domain computing arrangement scheme, i.e., p k,t =(α k,t ,o k,t ), α k,t ={α u,i (m)} k,t , o k,t ={o u,i (m)} k,t The initial cross-domain computing arrangement includes the ICPS domain selection result α for each large model training u,i (m) and the corresponding computing power resource allocation ratio ou,i (m); Each particle velocity corresponds to the adjustment direction and amplitude of the particle position, i.e. v k,t =(pr k,t ,v′ k,t ), pr k,t ={pr u,i (m)} k,t ,v′ k,t ={v′ u,i (m)} k,t The direction and amplitude of the particle position adjustment correspond to the probability of change of the ICPS domain selection result pr u,i (m) and the change in the computing power resource allocation ratio v′ u,i (m).

[0043] Step S103: Design a fitness function and evaluate the computational orchestration scheme. Specifically, the goal of the cross-domain computational orchestration algorithm based on particle swarm optimization is to find the best computational orchestration scheme to minimize the computational delay of model training. Therefore, the fitness function is related to the computational delay and is expressed as:

[0044]

[0045] Among them, c u,i Indicates the number of CPU cycles required to execute each bit of data. Indicates ICPS domain i u Computing resources (CPU frequency, i.e. the amount of data processed per unit time), o u,i (m) represents the proportion of computing resources allocated to model m, α u,i (m) = 1 means it is selected for training. The computational latency of model m will be determined by the slowest training ICPS domain. The smaller the fitness value, the better the particle's position and the more optimal the computational scheduling scheme.

[0046] Step S104: Calculate fitness and update the individual particle and global optimal position. Specifically, calculate and record the fitness values ​​of the particles in the population corresponding to the current iteration number, compare the current fitness value of each particle with the minimum fitness value in its historical iteration number, and determine the position of the particle with the smaller fitness value as the individual optimal position of the particle, expressed as The individual best position corresponds to the optimal arrangement scheme found by each particle in the historical search; the fitness values ​​of the individual best positions of all current particles are compared, and the individual best position of the particle with the smaller fitness value is determined as the global best position, which is expressed as b g =(α g ,o g ), the global optimal position corresponds to the optimal arrangement solution found by the entire population in the historical search.

[0047] Step S105: Update particle speed and position. Specifically, the particle swarm updates its speed and position by combining its individual best position and the global best position to find the optimal calculation arrangement. The particle speed update formula consists of two parts. The first part is expressed as:

[0048]

[0049] Among them, β1, β2 are weight coefficients, represents the position of particles in the population u,i (m) = the number of j, Record the arrangement scheme better and worse than particle p respectively k The second part is expressed as: Among them, ω, c1, c2 represent inertia weight, cognitive parameter and social parameter respectively. The elements in the matrix r1 and r2 are random numbers between [0, 1] to balance the local search and global exploration capabilities.

[0050] According to the position update formula, α k,t+1 ←pr k,t+1 , o k,t+1 =o k,t +v′ k,t+1 , adjust the position of each particle, modify the cross-domain computing orchestration plan, and evolve towards the optimization goal.

[0051] Step S106: Iterate and optimize to output the optimal computing orchestration solution. Specifically, repeat steps S104 and S105 until the iteration termination condition is reached, and output the global optimal position, i.e., the optimal cross-domain computing orchestration solution.

[0052] Step S107: Distribute the cross-domain computing orchestration plan, dynamically adjust, and continuously optimize. Specifically, the cloud controller distributes the optimal cross-domain computing orchestration plan to the terminal device layer for execution. Simultaneously, the cloud controller continuously monitors the resource status of each ICPS domain, restarts the cross-domain computing orchestration algorithm as needed, and dynamically adjusts the orchestration plan to ensure the continuous efficiency, reliability, and flexibility of the industrial large-scale model training process.

[0053] like Figure 3 As shown in Figure 1, the deterministic transmission scheduling algorithm based on proximal policy optimization (i.e., using a reinforcement learning algorithm to generate optimal traffic scheduling decisions) includes the following steps:

[0054] Step S201: Extract parameter flow and network status information. Specifically, the cloud controller is considered as an intelligent agent and extracts the parameter flow information and network status information to be scheduled. The parameter flow information includes the ID number, flow period, source address, destination address, optional transmission path, frame size, end-to-end delay requirement, jitter requirement, etc., which can be expressed as:

[0055]

[0056] in, is the total number of parameter flows; network status information includes link bandwidth B l , queue resource capacity of network communication equipment wait;

[0057] Step S202: Design the state space, action space and reward mechanism. Specifically,

[0058] The state space is defined as At time step t, the state Since the parameter flow is scheduled based on the transmission cycle, the queue resource margin is closely related to the transmission cycle (time), so a two-dimensional transmission cycle-queue resource margin matrix is ​​constructed. As part of the state space, it is expressed as:

[0059]

[0060] in, Represents the transmission period T L Internal node v m queue The action space is defined as At time step t, the action here, Refers to the optimal path selected from the optional transmission paths of the parameter flow, Determines the start time of sending the parameter stream. Determines which queue the parameter flow enters at the forwarding node.

[0061] The reward mechanism is used to evaluate action a t , to guide the agent to optimize its decision. Specifically, action a t The corresponding reward function r t for:

[0062]

[0063] in, Indicates that the parameter flow is in action a t The transmission delay under , is expressed as:

[0064]

[0065] in, Respectively represent the sending domain, network communication equipment v k The transmission period of the domain; Represents the flow f t In v kThe relative offset of the queue, Indicates the link delay between adjacent network communication devices. If the parameter flow scheduling is successful, a positive reward is given, and the smaller the delay, the greater the reward. If the scheduling fails, a negative constant penalty is given.

[0066] Step S203: Neural network training is performed to update parameters and output the optimal traffic scheduling decision. Specifically, the agent interacts with the environment and generates a series of t 、Action a t , reward r t , next state s t+1 For each time step t, the agent is based on the current state s t , the policy network outputs action a t , get immediate reward r after executing the action t And enter the next state s t+1 ; Value network receives state s t and s t+1 As input, output the estimated value V in two states φ (s t ),V φ (s t+1 ). During the training iteration, the policy network updates the parameters θ by maximizing the objective function J(θ), which is expressed as:

[0067]

[0068] Among them, ρ θ,t represents the ratio of the new and old strategies, ε is the exploration rate, and the advantage function It is used to measure the quality of the action relative to the average level, γ is the discount factor, λ is used to balance the deviation and variance, and T is the trajectory length. The policy network updates the parameter φ by minimizing the loss function L(φ), which is expressed as:

[0069]

[0070] Through the back-propagation mechanism of the neural network, the parameters of the policy network and the value network are continuously updated and optimized, enabling the intelligent agent to make the optimal traffic scheduling decision.

[0071] Step S204: Distribute traffic scheduling decisions, dynamically adjust, and continuously optimize. Specifically, traffic scheduling decisions are distributed to network communication devices in the network. These devices dynamically select the transmission path for the parameter stream, rationally schedule transmission times and queue allocation, and implement refined traffic control. Simultaneously, the cloud controller continuously monitors network status changes and new parameter stream requests, initiating a deterministic transmission scheduling algorithm based on proximal policy optimization as needed. Dynamically update traffic scheduling decisions to ensure deterministic parameter stream transmission throughout the entire lifecycle of industrial large-scale model training.

[0072] like Figure 4 As shown, the enhanced cyclic designated queuing and forwarding mechanism (i.e., multi-queue cyclic queuing and forwarding mechanism) includes the following steps:

[0073] Step S301: define the traffic scheduling paradigm of the forwarding device port queue; specifically, the forwarding device port is set Each queue has two working states: sending and receiving parameter flows, which are controlled by the sending gate and receiving gate respectively. A gate list is set to switch the queue working state by controlling the opening and closing of the sending gate and receiving gate. The time is divided into transmission cycles of equal length, with the transmission cycle T ζ The parameter stream is scheduled as the basic unit, where ζ represents the index number of the transmission period; in a certain transmission period T ζ There is only one queue in the sending state, represented by the value 0, and the rest Each queue is in the receiving state, represented by the value 1; the transmission cycle T ζ With queue Q m The corresponding relationship of the working status is expressed as:

[0074]

[0075] Among them, m represents the index number of the queue, let σ m For queue Q m The number of transmission cycles required to wait for switching to the sending state is expressed as:

[0076]

[0077] That is: in the transmission period T ζ Internal entry queue Q m The parameter stream (m≠ζ%N) will wait in the queue for σ m transmission cycles, in the transmission cycle T ζ+σm It is then sent to the next forwarding device.

[0078] Step S302: Measure the link delay and configure the start time of the transmission cycle of different forwarding devices. Specifically, for WADN scenarios, the network scale is large, the physical distance is long, and the link delay cannot be ignored. Measure the link delay between forwarding devices and set the start time of the same transmission cycle of adjacent forwarding devices to differ by the link delay to achieve staggered synchronization of the transmission cycles between forwarding devices and avoid data frame loss during transmission.

[0079] Step S303: Configure the forwarding device gating list based on the traffic scheduling decision and transmit the parameter stream. Specifically, the traffic scheduling decision determines the starting time for the parameter stream to be sent and which queue of the forwarding node in the transmission path it enters, thereby determining the opening and closing of the forwarding device port queue's send gate and receive gate within each transmission cycle. Within each transmission cycle, the forwarding device performs a send or receive operation on the parameter stream based on the configured forwarding device port queue status. For several forwarding device port queues within each transmission cycle, when the forwarding device port queue send gate is open, the parameter stream accumulated in the corresponding forwarding device port queue is forwarded to the corresponding downstream node on the parameter stream transmission path via the egress link; when the forwarding device port queue receive gate is open, the parameter stream from the upstream node on the parameter stream transmission path is cached in the corresponding forwarding device port queue, awaiting the arrival of the next transmission cycle. The forwarding device port queues refer to the several queues set for each port of the forwarding device.

[0080] The intelligent computing fusion network architecture provided in this embodiment is used to support the efficient construction of artificial intelligence-generated content (AIGC) services in industrial cyber-physical systems (ICPS). The intelligent computing fusion network architecture includes: a terminal device layer, an edge service layer, a core network layer, and a cloud control layer. The terminal device layer trains and updates local models based on local data sample sets. The edge service layer is used to connect terminal devices to the core network layer and perform edge aggregation of local model parameters. The core network layer deploys a multi-queue round-robin queuing forwarding mechanism on forwarding devices such as routers to provide bounded latency, jitter, and packet loss rate for parameter stream transmission, thereby achieving reliable transmission scheduling in the wide area network. The cloud control layer is used to achieve global aggregation of model parameters and dynamic adaptation of computing network resources. The new network architecture designed in this embodiment supports AIGC model training and efficiently allocates distributed and diversified computing resources and network resources, promoting the implementation of AIGC services in the industrial production field, optimizing the overall latency of AIGC model training, and achieving cross-domain computing collaboration of terminal devices and end-to-end reliable transmission scheduling of model parameters.

[0081] Example 2

[0082] This embodiment further provides a computing network resource adaptation and reliable transmission scheduling method, which applies the intelligent computing fusion network architecture described in Example 1. The computing network resource adaptation and reliable transmission scheduling method includes:

[0083] SA: The cloud control layer uses a particle swarm algorithm to generate an optimal cross-domain orchestration solution based on the computing resources of each industrial cyber-physical system domain and the training requirements of the AIGC model. The optimal cross-domain orchestration solution includes the industrial cyber-physical system domain selection results and the computing resource allocation ratio of the corresponding industrial cyber-physical system domain.

[0084] SB: The terminal device layer determines the target industrial cyber-physical system domain according to the optimal cross-domain orchestration solution. The target industrial cyber-physical system domain trains a local model based on the local sample data set and allocation parameters to obtain local model training parameters.

[0085] SC: The edge service layer performs edge aggregation processing on the local model training parameters to obtain edge model parameters;

[0086] SD: The cloud control layer performs global aggregation processing on the edge model parameters to obtain global model parameters.

[0087] The computing network resource adaptation and reliable transmission scheduling method further includes:

[0088] The cloud control layer uses a deep reinforcement learning algorithm to generate an optimal traffic scheduling decision based on the parameter flow information to be scheduled and the network status information, and determines the parameter flow transmission path based on the optimal traffic scheduling decision and configures a gating list for each network communication device on the parameter flow transmission path, wherein the parameter flow information includes an ID number, a flow period, a source address, a destination address, an optional transmission path, a frame size, an end-to-end delay requirement, and a jitter requirement; the network status information includes the link bandwidth and the port queue resource capacity of the network communication device; the network communication device includes a forwarding device in the core network, a switch in the terminal device layer, and an access device in the edge service layer;

[0089] Each network communication device on the parameter stream transmission path performs parameter stream transmission according to the corresponding gating list.

[0090] When the network communication device on the parameter stream transmission path is a forwarding device in the core network layer, and the gating list is a forwarding device gating list, the forwarding device in the core network layer performs parameter stream transmission according to the forwarding device gating list, specifically including:

[0091] For several forwarding device port queues in each transmission cycle, when the forwarding device port queue sending gate is in the open state, the parameter flow accumulated in the corresponding forwarding device port queue is forwarded to the corresponding downstream node on the parameter flow transmission path through the egress link; when the forwarding device port queue receiving gate is in the open state, the parameter flow from the upstream node on the parameter flow transmission path is cached in the corresponding forwarding device port queue, waiting for the arrival of the next transmission cycle, wherein the forwarding device port queue refers to several queues set at the port of each forwarding device.

[0092] In this embodiment, the process of using the particle swarm algorithm to generate the optimal cross-domain orchestration solution, the process of using the reinforcement learning algorithm to generate the optimal traffic scheduling decision, and the process of using the multi-queue circular queuing forwarding mechanism to transmit parameter streams can be referred to the specific execution process of the cloud control layer in Example 1, and will not be repeated here.

[0093] The computing network resource adaptation and reliable transmission scheduling method provided in this embodiment serves the intelligent computing fusion network architecture described in Example 1. The intelligent computing fusion network architecture fully utilizes computing network resources through the coordinated cooperation among the four layers of the terminal device layer, edge service layer, core network layer and terminal control layer, realizes cross-domain computing collaboration and end-to-end reliable transmission scheduling of parameter flows, optimizes the computing delay and transmission delay of AIGC model training, and promotes the efficient construction of AIGC models in ICPS; the cross-domain computing orchestration algorithm based on particle swarm optimization optimizes the computing delay by selecting suitable computing devices and allocating computing power resources for AIGC model training, ensuring the efficient operation of training. The deterministic transmission scheduling algorithm based on proximal policy optimization optimizes the transmission delay of parameter flows and accelerates the aggregation process of parameter flows by reasonably planning queue (time) resources for parameter flow transmission. The multi-queue circular queuing forwarding mechanism ensures that the parameter flows can be transmitted according to the traffic scheduling decision by defining the traffic scheduling paradigm in the underlying forwarding device, thereby realizing end-to-end reliable transmission scheduling of parameter flows.

[0094] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0095] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. An intelligent computing fusion network system, characterized in that: The intelligent computing fusion network system includes: The terminal device layer includes a plurality of industrial cyber-physical system domains distributed in different regions or factories, each of which includes a computing device, each of which is used to train a local model based on a local sample data set and target parameters to obtain local model training parameters, wherein the target parameters are the AIGC model parameters that need to be updated and are allocated by the cloud control layer to each of the industrial cyber-physical system domains; An edge service layer, connected to the terminal device layer, for performing edge aggregation processing on the local model training parameters to obtain edge model parameters; The cloud control layer is connected to the edge service layer through the core network layer and is used to perform global aggregation processing on the edge model parameters to obtain global model parameters; each forwarding device in the core network layer transmits parameter streams based on a multi-queue round-robin forwarding mechanism, wherein the multi-queue round-robin forwarding mechanism refers to controlling the working status of several forwarding device port queues within each transmission cycle through a forwarding device gating list; The cloud control layer is also used to generate the optimal cross-domain orchestration solution using a particle swarm algorithm and the optimal traffic scheduling decision using a deep reinforcement learning algorithm. Specifically: According to the computing power resources of each industrial cyber-physical system domain and the training requirements of the AIGC model, a particle swarm algorithm is used to generate an optimal cross-domain orchestration plan, wherein the optimal cross-domain orchestration plan includes the industrial cyber-physical system domain selection result and the computing power resource allocation ratio of the corresponding industrial cyber-physical system domain; according to the parameter flow information to be scheduled and the network status information, a deep reinforcement learning algorithm is used to generate an optimal traffic scheduling decision, and the parameter flow transmission path is determined according to the optimal traffic scheduling decision and the gating list of each network communication device on the parameter flow transmission path is configured, wherein the parameter flow information includes an ID number, a flow period, a source address, a destination address, an optional transmission path, a frame size, an end-to-end delay requirement and a jitter requirement, the network status information includes the link bandwidth and the port queue resource capacity of the network communication device, the network communication device includes the forwarding device in the core network, the switch in the terminal device layer and the access device in the edge service layer, and the gating list includes the above-mentioned switch gating list control and the forwarding device gating list.

2. The intelligent computing fusion network system according to claim 1, characterized in that: The terminal device layer is connected to the edge service layer through a switch, wherein the switch is used for a time-sensitive network transmission parameter stream based on a cyclic queuing forwarding mechanism. The cyclic queuing forwarding mechanism refers to controlling the switching of the working states of two switch port queues in a preset cyclic time slot through a switch gating list. The working states include receiving parameter streams and sending parameter streams. The switch port queues refer to two queues set at the ports of each switch. The switch gating list represents the opening and closing states of the switch queue sending gate and the switch queue receiving gate in each preset cyclic time slot. When the switch queue sending gate is open, the switch port queue sends the parameter stream, and when the switch queue receiving gate is open, the switch port queue receives the parameter stream.

3. The intelligent computing fusion network system according to claim 1, characterized in that: The edge service layer includes several device groups, each of which includes an access device and an edge server; The access device is used to connect a terminal device to the core network layer, wherein the computing device is a type of terminal device; The edge server is used to perform edge aggregation processing on the local model training parameters corresponding to the industrial information-physical system domain within a preset coverage area to obtain the edge model parameters.

4. The intelligent computing fusion network system according to claim 1, characterized in that: The forwarding device port queue refers to several queues set at the port of each forwarding device, and the working status includes receiving parameter flow and sending parameter flow. In each transmission cycle, only one forwarding device port queue has a working status of sending parameter flow, and the working status of the remaining forwarding device port queues is receiving parameter flow. The forwarding device gating list represents the opening and closing status of the forwarding device port queue sending gate and the forwarding device port queue receiving gate in each transmission cycle. The forwarding device port queue sending gate is open for the forwarding device port queue sending parameter flow, and the forwarding device port queue receiving gate is open for the forwarding device port queue receiving parameter flow.

5. The intelligent computing fusion network system according to claim 4, characterized in that: The time difference between the start times of the same transmission cycle of two adjacent forwarding devices is equal to the link delay between the forwarding devices.

6. A method for network resource adaptation and reliable transmission scheduling, characterized in that: The computing network resource adaptation and reliable transmission scheduling method applies the intelligent computing fusion network system described in any one of claims 1 to 5, and the computing network resource adaptation and reliable transmission scheduling method includes: The cloud control layer uses a particle swarm algorithm to generate an optimal cross-domain orchestration solution based on the computing resources of each industrial cyber-physical system domain and the training requirements of the AIGC model. The optimal cross-domain orchestration solution includes the industrial cyber-physical system domain selection result and the computing resource allocation ratio of the corresponding industrial cyber-physical system domain. The terminal device layer determines the target industrial cyber-physical system domain according to the optimal cross-domain orchestration solution, and the target industrial cyber-physical system domain trains a local model according to the local sample data set and the allocation parameters to obtain local model training parameters; The edge service layer performs edge aggregation processing on the local model training parameters to obtain edge model parameters; The cloud control layer performs global aggregation processing on the edge model parameters to obtain global model parameters.

7. The method for network resource adaptation and reliable transmission scheduling according to claim 6, characterized in that: The computing network resource adaptation and reliable transmission scheduling method further includes: The cloud control layer uses a deep reinforcement learning algorithm to generate an optimal traffic scheduling decision based on the parameter flow information to be scheduled and the network status information, and determines the parameter flow transmission path based on the optimal traffic scheduling decision and configures a gating list for each network communication device on the parameter flow transmission path, wherein the parameter flow information includes an ID number, a flow period, a source address, a destination address, an optional transmission path, a frame size, an end-to-end delay requirement, and a jitter requirement; the network status information includes the link bandwidth and the port queue resource capacity of the network communication device; the network communication device includes a forwarding device in the core network, a switch in the terminal device layer, and an access device in the edge service layer; Each network communication device on the parameter stream transmission path performs parameter stream transmission according to the corresponding gating list.

8. The method for network resource adaptation and reliable transmission scheduling according to claim 7, characterized in that: When the network communication device on the parameter stream transmission path is a forwarding device in the core network layer, and the gating list is a forwarding device gating list, the forwarding device in the core network layer performs parameter stream transmission according to the forwarding device gating list, specifically including: For several forwarding device port queues in each transmission cycle, when the forwarding device port queue sending gate is in the open state, the parameter flow accumulated in the corresponding forwarding device port queue is forwarded to the corresponding downstream node on the parameter flow transmission path through the egress link; when the forwarding device port queue receiving gate is in the open state, the parameter flow from the upstream node on the parameter flow transmission path is cached in the corresponding forwarding device port queue, waiting for the arrival of the next transmission cycle, wherein the forwarding device port queue refers to several queues set at the port of each forwarding device.

Citation Information

Patent Citations

  • Industrial internet network architecture based on intelligent fusion of general inductance calculation and virtual controller

    CN116962409A

  • Deterministic transmission scheduling system oriented to calculation network integration and working method of deterministic transmission scheduling system

    CN119052337A