Data processing method and device, equipment and medium
By employing local state awareness and on-path update mechanisms, the problem of slow parameter convergence in centralized controllers in large-scale networks is solved, achieving low-latency, high-reliability computing network resource adaptation and improving the dynamic adaptability of computing power services.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies for large-scale distributed computing networks, centralized controllers struggle to converge scheduling parameters within a limited timeframe, failing to adapt to the dynamic demands of computing power services, resulting in low resource matching accuracy and excessive latency.
By adopting a local state awareness and local policy model, and through the in-path update mechanism of source and receiver nodes, the scheduling policy is dynamically adjusted to achieve frame-level response and fine-grained resource adaptation.
It simplifies scheduling complexity, improves resource utilization and the parallel task processing capability of computing clusters, and provides low-latency, highly reliable integrated computing and network services.
Smart Images

Figure CN121887800A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a data processing method, apparatus, device and medium. Background Technology
[0002] With the development of computing networks, optical computing networks have become the core carrier for large-scale distributed computing tasks. To achieve fine-grained scheduling of computing resources, fine-grained optical transport networks (fgOTN) introduce a multi-layered mapping mechanism: from fine-grained units of the computing network to optical payload units (OPUs), then to optical data units (ODUk), and finally to optical transport units (OTUk). While this mechanism provides bandwidth flexibility, it also necessitates joint decision-making across multiple dimensions for each service scheduling operation.
[0003] In related technologies, a centralized intelligent algorithm deployed on the controller is typically used to solve for the aforementioned scheduling parameters. The controller receives status information reported by each node, makes centralized decisions based on the global state, and then distributes the generated scheduling parameters to the corresponding nodes for execution. However, as the network scale increases and the mapping layers increase, the algorithm struggles to converge within a limited time, making it difficult to adapt to the dynamic demands of computing power services. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a data processing method, apparatus, device and medium.
[0005] To achieve the above objectives, this application provides a data processing method applied to a source node, comprising: calculating a scheduling strategy based on local state information perceived by the source node and a received local policy model; the local state information including service characteristics, network link status, and computing node status; the scheduling strategy including target payload block, target optical path data unit, target routing information, and target computing node; the local policy model being obtained by a controller based on a global policy model; processing frame information based on service data, the local state information, and the scheduling strategy; and updating the scheduling strategy along the path by the receiving node during transmission based on the perceived local state information and the corresponding local policy model.
[0006] Optionally, the scheduling strategy is calculated based on the local state information perceived by the source node and the received local policy model, including: obtaining the network load entropy based on the link utilization in the network link state, the network load entropy being used to measure link load balance; obtaining the computing load entropy based on the computing load degree in the computing node state, the computing load entropy being used to measure computing node balance; obtaining the service fusion entropy based on the service type proportion in the service characteristics, the service fusion entropy being used to measure service compatibility within the same optical path data unit; obtaining the current coupling entropy based on the network load entropy, the computing load entropy, and the service fusion entropy; and obtaining the scheduling strategy based on the current coupling entropy and the local policy model.
[0007] Optionally, the scheduling strategy is obtained based on the current coupling entropy and the local policy model, including: constructing a current computing network state based on the current coupling entropy, the bandwidth requirement in the network link state, and the computing power requirement in the computing power node state; calculating a similarity between the current computing network state and the historical computing network state; if the similarity is greater than a state threshold, then obtaining a scheduling strategy corresponding to the current computing network state based on the historical scheduling strategy corresponding to the historical computing network state; otherwise, inputting the current computing network state into the local policy model to generate a scheduling strategy for the current computing network state with the goal of maximizing the scheduling reward constructed based on the change in coupling entropy, the comprehensive utilization rate of computing network resources, and the transmission delay; the computing network resource utilization rate is obtained based on the utilization rate of payload block resources, the utilization rate of optical path data unit resources, and the utilization rate of computing power nodes.
[0008] Optionally, the frame information is obtained by processing the service data, the local state information, and the scheduling strategy, including: determining the target optical path load unit based on the service characteristics in the local state information; synchronous services in the same service group are located in the same target optical path load unit, and the synchronous services are determined by the service characteristics; mapping the service data and the local state information to the target load block; mapping the target load block to the target optical path load unit; the number of load blocks in the target optical path load unit does not exceed a number threshold; mapping the target optical path load units of the same synchronous service to the same target optical path data unit; the target load block occupies a continuous time slot position in the target optical path data unit; and mapping the target optical path data unit to the target optical path transmission unit to obtain the frame information.
[0009] Another embodiment of this application provides a data processing method applied to an intermediate node, comprising: receiving local state information perceived by the intermediate node, frame information sent by the previous node, and a scheduling strategy; the previous node being a source node or an intermediate node; the frame information including service data and local state information of the upstream node; the local state information including service characteristics, network link status, and computing power node status; the scheduling strategy of the previous node including target payload block, target optical path data unit, target routing information, and target computing power node, which is obtained by dynamically updating the scheduling strategy of the source node along the path during transmission based on the local state information of the upstream node and the corresponding local strategy model; decapsulating the frame information according to the scheduling strategy of the previous node to extract the service data and the local state information of the upstream node; calculating a new scheduling strategy based on the local state information of the upstream node, the local state information perceived by the intermediate node, and the received local strategy model; and processing the service data, the local state information of the intermediate node, and the new scheduling strategy to obtain new frame information.
[0010] Another embodiment of this application provides a data processing method applied to a destination node, comprising: receiving frame information and a corresponding scheduling strategy from a previous node; the previous node being a source node or an intermediate node; the frame information including service data and local state information of an upstream node; the scheduling strategy including a target payload block, a target optical path data unit, target routing information, and a target computing power node; the scheduling strategy of the previous node being dynamically updated along the path of the source node's scheduling strategy during transmission based on the local state information of the upstream node and the corresponding local strategy model; decapsulating the frame information according to the scheduling strategy of the previous node to extract the service data; and sending the service data to the target computing power node for processing to obtain computing power execution results.
[0011] Optionally, the method further includes: receiving the computing node status reported by the target computing node; obtaining the local state information of the target node based on the computing node status; and reporting the local state information of the target node to the controller for updating the global policy model, wherein the updated global policy model is used to update the local policy model.
[0012] Optionally, the method further includes: monitoring the execution status of the target computing node on the business data; if the business data has not been processed, then making the target node a new source node, and recalculating and generating scheduling policies and frame information based on its currently perceived local state information and local policy model, so as to unload the remaining business data to subsequent computing nodes for execution.
[0013] This application provides a data processing apparatus, including: a receiving module for receiving local state information and a local policy model; and a calculation module for calculating a scheduling policy based on the local state information and the local policy model. The local state information includes service characteristics, network link status, and computing node status. The scheduling policy includes target payload blocks, target optical path data units, target routing information, and target computing nodes. The local policy model is obtained by a controller based on a global policy model. Frame information is obtained by processing service data, the local state information, and the scheduling policy. The scheduling policy is updated along the path by the receiving node during transmission based on its own perceived local state information and the corresponding local policy model.
[0014] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method.
[0015] As can be seen from the above, this application proposes a data processing method. At the response time level, status information is transmitted and dynamically updated in real-time along the transmission path at the frame-level granularity. This mechanism breaks the vertical interaction loop of "node reporting - central decision-making - instruction issuance" in related technologies, avoiding the processing load of centralized controllers when handling high-concurrency services, and significantly reducing scheduling latency from the periodic level of related technologies to frame-level response. At the resource regulation level, by establishing a dynamic mapping mechanism between payload blocks and ODUs, fine-grained on-demand adaptation of computing power services and optical layer resources can be achieved. While ensuring that service characteristics are met, it can significantly improve the comprehensive utilization rate of computing network resources and the parallel task processing capability of computing power clusters. At the policy evolution level, receiving nodes can perform path-by-path correction of scheduling policies based on real-time perceived local states and corresponding local policy models, which can solve the problems of information asymmetry and inaccurate decision-making caused by the lag of status information in large-scale networking environments. In summary, the embodiments of this application simplify the complexity of computing network scheduling while providing low-latency, high-reliability and highly deterministic carrying capacity for integrated computing network services, and can adapt to the dynamic needs of computing power services. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram illustrating an application scenario of a data processing method according to an embodiment of this application; Figure 2 This is a flowchart illustrating a data processing method applied to a source node according to an embodiment of this application. Figure 3 This is a schematic diagram of the frame structure and multi-layer mapping relationship of an embodiment of this application; Figure 4 This is a flowchart illustrating a data processing method applied to an intermediate node according to an embodiment of this application. Figure 5 This is a flowchart illustrating a data processing method applied to a destination node according to an embodiment of this application. Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0020] With the rise of computationally intensive applications such as large-scale AI model training and real-time intelligent inference, Computing Power Networks (CPNs) have become an information infrastructure that integrates computing, storage, and network resources. In this architecture, the Computing-Aware Optical Network (CPN) serves as the physical platform, undertaking cross-domain collaboration and data transmission for large-scale distributed computing tasks.
[0021] To meet the diverse needs of computing services for bandwidth granularity, scheduling flexibility, and quality of service, coarse-grained optical transport networks with wavelength as the smallest scheduling unit are insufficient. Therefore, a fine-grained optical transport network (fgOTN) architecture is proposed. This fgOTN architecture introduces a multi-layer mapping mechanism: the fine-grained unit of the computing optical network is mapped to the optical path payload unit (OPU); subsequently, the OPU is encapsulated into the optical path data unit (ODUk); then, the ODUk is further mapped to the optical path transmission unit (OTUk); finally, the OTUk signal is modulated and loaded onto the specified wavelength channel, completing cross-domain, low-latency transparent transmission at the optical layer.
[0022] In optical computing networks, a fine-grained unit (FPU) refers to the smallest schedulable service unit (SSU) with specific computing power requirements and network service quality (SQV) in the optical computing network. This unit encapsulates both computing and communication attributes. The Optical Processing Unit (OPU) is one of the key logical layers in the optical transport frame structure, primarily responsible for adapting, mapping, and adjusting the rate of customer service signals, providing standardized payload encapsulation for the upper-layer Optical Distribution Unit (ODUk). The Optical Distribution Unit (ODUk) is also a key logical layer in the optical transport frame structure, located above the OPU and below the OTU. It undertakes key functions such as end-to-end data transmission, performance monitoring, fault isolation, and sub-wavelength multiplexing, forming the foundation for highly reliable, maintainable, and flexibly schedulable optical transmission. The Optical Transmission Unit (OTUk) is the outermost digital frame structure in the optical transport frame structure, located above the ODUk, and directly transmits to the physical optical layer. It is a key carrier for achieving long-distance, highly reliable, and maintainable optical transmission, undertaking core functions such as forward error correction (FEC), segment-level monitoring, and optical channel adaptation. In ODUk and OTUk, 'k' represents the rate level identifier, with different levels corresponding to different nominal bandwidths and carrying capacities. In addition, ODUk also supports the flexible container ODUflex, whose rate can be dynamically adjusted according to the customer's business signal requirements.
[0023] While this multi-layered mapping mechanism significantly improves resource utilization efficiency and service adaptability, it struggles to meet the dynamic demands of computing power services. A single computing power service network access scheduling requires information such as capacity matching of fine-grained units in the optical network, ODUk type adaptation, wavelength routing, and computing node allocation. Related technical solutions rely on a centralized intelligent scheduling algorithm deployed on the controller to uniformly process these coupled variables. However, as network scale expands, mapping layers deepen, and computing power services exhibit dynamic characteristics such as high concurrency, strong bursts, and short lifecycles, this approach suffers from the following drawbacks: the number of decision variables increases polynomially or even exponentially with the number of services, nodes, and ODUk types, making it difficult for centralized algorithms to converge quickly in large-scale scenarios. All scheduling requests converge on a single controller, easily creating performance bottlenecks and failing to effectively handle non-steady-state loads such as bursty traffic and random requests from edge inference. Periodic delays in the status updates of the optical layer (e.g., wavelength occupancy) and computing layer (e.g., GPU utilization, queue latency) lead to scheduling decisions based on outdated information, reducing resource matching accuracy.
[0024] Furthermore, the scheduling mechanisms of related technologies typically optimize resources in isolation, focusing solely on maximizing network link bandwidth utilization or load balancing of computing nodes, while ignoring the heterogeneity of computing services in terms of computational characteristics and network service quality. This blind scheduling strategy leads to services of different service levels being indiscriminately reused in the same ODU container for transmission. For example, if low-latency inference tasks and high-bandwidth data transmission share the same ODU channel, although it improves link resource utilization at the statistical reuse level, uncontrollable queuing latency and jitter can occur due to queue contention, lack of scheduling priority, or buffer congestion, leading to asynchronous latency and degraded computing efficiency.
[0025] Furthermore, optical network scheduling in related technologies generally adopts a "fill-first" strategy, which prioritizes mapping services to established but not fully loaded high-order ODU containers to improve link resource utilization. However, to adapt to existing fixed-rate containers, fine-grained services are forced to "align up" to coarse-grained ODUs that far exceed their actual needs, resulting in significant bandwidth fragmentation and resource waste. In addition, fine-grained services often need to detour rather than take the shortest path when reusing high-order ODU containers, leading to an increase in optical layer hops and a longer transmission distance. This, in turn, adds to the propagation delay and node processing delay, causing the end-to-end latency to exceed the requirements of low-latency computing tasks.
[0026] Based on this, this application proposes a data processing method. At the response time level, status information is transmitted and dynamically updated in real-time along the transmission path at the frame-level granularity. This mechanism breaks the vertical interaction loop of "node reporting - central decision-making - instruction issuance" in related technologies, avoiding the processing load of centralized controllers when handling high-concurrency services, and significantly reducing scheduling latency from the periodic level of related technologies to frame-level response. At the resource regulation level, by establishing a dynamic mapping mechanism between payload blocks and ODUs, fine-grained on-demand adaptation of computing power services and optical layer resources can be achieved. While ensuring that service characteristics are met, the overall utilization rate of computing network resources and the parallel task processing capability of computing power clusters can be significantly improved. At the policy evolution level, receiving nodes can perform path-by-path correction of scheduling policies based on real-time perceived local states and corresponding local policy models, solving the problems of information asymmetry and inaccurate decision-making caused by the lag of status information in large-scale networking environments. In summary, the embodiments of this application simplify the complexity of computing network scheduling while providing low-latency, high-reliability and highly deterministic carrying capacity for integrated computing network services, and can adapt to the dynamic needs of computing power services.
[0027] Furthermore, by coordinating constraints on the spatial adjacency (time slot continuity) and logical normalization (same OPU / ODUk mapping) of payload blocks during the mapping process, physical-level isolation between services can be achieved, avoiding queue competition and queuing delays for services with different characteristics within the shared channel. This solves the problem of asynchronous computing task timing caused by transmission fluctuations, thereby avoiding computational efficiency degradation and providing physical channel guarantees for fine-grained computing power collaboration across nodes.
[0028] Furthermore, by leveraging local state awareness, precise matching between payload blocks and service requirements is achieved, eliminating bandwidth fragmentation and improving network resource utilization. The target routing information in the scheduling strategy supports path planning with latency as the core constraint, breaking the limitation of detours caused by improving reuse rate. By shortening transmission distance and reducing forwarding hops, physical link latency is significantly reduced. Combined with the path-as-you-go policy update mechanism, it can ensure that fine-grained computing services are always transmitted along the path with optimal latency in a dynamically changing topology environment.
[0029] refer to Figure 1 This is a schematic diagram illustrating an application scenario of the data processing method provided in this application embodiment. The application scenario includes a controller 101, a source node 102, a receiving node, and a computing node 104. The controller 101, source node 102, receiving node, and computing node 104 can all be connected via wired fiber optic links or wireless communication networks to jointly construct a computing optical network.
[0030] Controller 101 can be a standalone physical server, a cluster or distributed system consisting of multiple physical servers, or a software-defined network controller deployed in the cloud. Controller 101 can be responsible for the orchestration of global resources and the distribution of policies. For example, it can distribute local policy models to source node 102 and receiving node.
[0031] The source node 102 and the receiving node can be edge nodes or aggregation nodes in the optical computing network, including but not limited to optical transceivers, optical transmission equipment, edge computing gateways or other electronic devices with photoelectric conversion and data processing capabilities.
[0032] The source node 102 can be used to sense local state information, receive the local policy model sent by the controller 101, calculate and generate a scheduling policy based on the local policy model and the local state information it senses, and process the frame information based on the service data, local state information and scheduling policy.
[0033] The receiving node can be either an intermediate node 103a or a destination node 103b. In this embodiment, the receiving node may have data forwarding or receiving functions, as well as path-to-path update capabilities. That is, the receiving node can capture local real-time fluctuations in the computing network state based on its own perceived real-time local state information, and modify the scheduling strategy according to the local state information and the corresponding local strategy model, realizing the leap from static preset at the source end to dynamic adaptive scheduling decision-making along the path. In this embodiment, the source node, intermediate node, and / or destination node may correspond to the same node, which performs different functions in different links. When the receiving node is an intermediate node 103a, the intermediate node 103a can receive its own perceived local state information, frame information sent by the previous node, and scheduling policy; the previous node can be a source node or an intermediate node. It decapsulates the frame information according to the previous node's scheduling policy to extract service data and the local state information of the upstream node. It calculates a new scheduling policy based on the local state information of the upstream node, the local state information perceived by the intermediate node, and the received local policy model. It then processes the service data, the intermediate node's local state information, and the new scheduling policy to obtain new frame information. When the receiving node is a destination node 103b, the destination node 103b can receive the frame information and corresponding scheduling policy of the previous node; the previous node can be a source node or an intermediate node. It decapsulates the frame information according to the previous node's scheduling policy to extract service data. It then sends the service data to the computing power node 104 for processing to obtain the computing power execution result. During frame information transmission, the intermediate node and the destination node can change according to the scheduling policy.
[0034] The computing node 104 can refer to the infrastructure providing computing resources in the computing power optical network, which can be a data center, edge computing center, high-performance computing cluster, or server carrying artificial intelligence tasks. The source node 102 and the receiving node can be connected to at least one computing node 104, either separately or simultaneously. In this embodiment, the computing node 104 can have dual functions of task collaborative processing and computing power status sensing: on the one hand, it can be responsible for carrying and processing various computing power service tasks offloaded from the destination node; on the other hand, the computing node 104 can dynamically feed back its internal computing node status to the associated source node 102 and / or receiving node. This status information can be used as input for in-path scheduling decisions, supporting the receiving node to dynamically update the global policy model of the controller according to the current computing node status, thereby updating the local policy model issued by the controller to the node.
[0035] The data processing method of this application can be applied to scenarios such as East-West data processing, industrial internet, autonomous driving, telemedicine, cross-regional AI model training, and ultra-large-scale scientific computing. By deeply integrating computing power perception with optical network scheduling, it provides a low-complexity, low-latency, highly reliable, and strongly deterministic integrated computing and network service system.
[0036] The following is combined Figure 1 The application scenarios described above illustrate the data processing method according to exemplary embodiments of this application. It should be noted that the above application scenarios are merely shown to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way in this regard. Rather, the embodiments of this application can be applied to any applicable scenario.
[0037] refer to Figure 2 The flowchart illustrates the data processing method provided in this application embodiment, which, when applied to the source node, may include: S1a. The scheduling strategy is calculated based on the local state information perceived by the source node and the received local policy model. The local state information includes service characteristics, network link status and computing node status. The scheduling strategy includes target payload block, target optical path data unit, target routing information and target computing node. The scheduling strategy is obtained by the controller based on the global policy model.
[0038] S2a. Frame information is obtained by processing business data, local state information and scheduling strategy. The scheduling strategy is updated by the receiving node during transmission based on the perceived local state information and the corresponding local strategy model.
[0039] In this embodiment, local state information refers to multi-dimensional environmental parameters acquired in real time by the source or receiving node through its state awareness module during network operation. Local state information is highly dynamic and can reflect the real-time load and demand profile of resources within the node's neighborhood. Local state information may include, but is not limited to, service characteristics, network link status, and computing node status.
[0040] Business characteristics can be used to describe the features of the business data to be processed and can be collected through the business access gateway. In one implementation, the sampling period is ≤100ms. These business characteristics may include, but are not limited to: bandwidth requirements, computing power requirements, business type, priority, and the business group to which they belong. For example, bandwidth requirements can be bandwidth time slot requirements, computing power requirements can be floating-point operations per second (FLOPS) requirements, and business types can be categorized as generative, synchronous, and independent. For instance, the bandwidth requirements in the Artificial Intelligence (AI) inference service G1... For 200Mbps, computing power requirement 500 FLOPS, business type Synchronous, priority It is high.
[0041] Network link status can be used to reflect the carrying capacity of physically adjacent links of nodes and can be acquired by optical power sensors. This network link status may include, but is not limited to: wavelength utilization, payload block resource utilization, OPU resource utilization, and optical path data unit resource utilization (i.e., ODUk resource utilization) of each link. Optionally, the resource utilization among payload block resource utilization, OPU resource utilization, and ODUk resource utilization may refer to time slot occupancy rate.
[0042] The computing node status can be used to describe the operating status of the computing nodes connected to the node, and can be collected by the performance monitoring unit of the computing node. The computing node status may include, but is not limited to, computing node utilization, for example, the real-time FLOPS utilization of the computing node, storage resource utilization, etc.
[0043] The local policy model can be dynamically updated by the controller based on the global policy model, which in turn is updated dynamically based on the state information reported by multiple nodes and the corresponding local policy model parameters. This local policy model can be distributed to each node by the controller. The local state information perceived by the source node is processed based on the local policy model distributed by the controller to the source node to obtain the scheduling policy. During transmission, the receiving node updates the scheduling policy along the path based on the perceived local state information and the corresponding local policy model. The local policy models distributed by the controller to different nodes may differ. In one implementation, the global policy model or the local policy model can be one or more combinations of heuristic algorithms, reinforcement learning models, or policy evaluation models.
[0044] In this embodiment, the scheduling strategy can be used to guide the parameters of frame information generation, and can be obtained with the goal of minimizing the overall utilization rate of computing network resources. This scheduling strategy may include, but is not limited to: target payload blocks, target optical path data units, target routing information, and target computing nodes.
[0045] A target payload block refers to a payload block (PB) determined by the local policy network for generating frame information. A frame can include one or more target payload blocks, and the sizes of different payload blocks can be the same or different. A fine-grained unit of a computing power optical network can be composed of one or more target payload blocks combined through a mapping scheme. By flexibly selecting payload blocks in space and time, fine-grained service slicing isolation can be achieved.
[0046] The target optical path data unit (ODU) can refer to the ODU used to generate frame information, as determined by the local policy network, and the target payload block is mapped to this target ODU. In one implementation, payload blocks of the same service can be mapped to the same ODUk to achieve logical hard isolation between services.
[0047] Destination routing information may refer to the transmission path planned for frame information as determined by the local policy network. This destination routing information may include the path from the source node to the destination node.
[0048] The target computing power node can refer to the computing power node allocated to the target node by the local policy network to offload business tasks.
[0049] After obtaining the scheduling policy, the source node can process the service data, local state information, and scheduling policy to obtain frame information. This frame information may include service data and local state information.
[0050] As-along update refers to updating the scheduling information along the transmission path. During transmission, the receiving node does not simply passively forward the frame; instead, it dynamically modifies the original scheduling decision based on its own perceived local state information and local policy model to obtain a new scheduling policy. The receiving node obtains the new frame information based on the new scheduling policy and its own local state information, and then forwards the new frame information to the next receiving node according to the target routing information in the new scheduling policy, until it reaches the destination node. The target routing information and corresponding destination node planned by the new scheduling policy may be the same as or different from those planned by the original scheduling policy. This as-along update mechanism ensures that the scheduling decision can adaptively evolve according to real-time fluctuations in the transmission path, avoiding latency in the backhaul controller, thus avoiding invalid waiting and synchronization losses during the computation process. Therefore, without increasing physical computing hardware, it improves the overall computing power efficiency of the entire network through extreme network performance optimization, achieving "computing with the network."
[0051] In terms of response timeliness, this application's embodiments transmit and dynamically update status information in real time along the transmission path at the frame-level granularity. This mechanism breaks the vertical interaction loop of "node reporting - central decision-making - instruction issuance" in related technologies, avoiding the processing load of centralized controllers when handling high-concurrency services, and significantly reducing scheduling latency from the periodic level of related technologies to frame-level response. At the resource regulation level, by establishing a dynamic mapping mechanism between payload blocks and ODUs, fine-grained on-demand adaptation of computing power services and optical layer resources can be achieved. While ensuring that service characteristics are met, it can significantly improve the comprehensive utilization rate of computing network resources and the parallel task processing capability of computing power clusters. At the policy evolution level, receiving nodes can perform path-by-path correction of scheduling policies based on real-time perceived local states and corresponding local policy models, solving the problems of information asymmetry and inaccurate decision-making caused by the lag of status information in large-scale networking environments. In summary, this application's embodiments simplify the complexity of computing network scheduling while providing low-latency, highly reliable, and highly deterministic carrying capacity for integrated computing network services, adapting to the dynamic needs of computing power services.
[0052] In an optional embodiment, the scheduling policy is calculated based on the local state information perceived by the source node and the received local policy model, which may include: S100. Based on the link utilization in the network link status, obtain the network load entropy, which is used to measure the link load balance.
[0053] In one implementation, the expression for calculating network load entropy can be:
[0054] in, It can refer to network load entropy. "E" can refer to the utilization rate of link "e", where "E" can refer to the set of network links. For example, if the utilization rates of the four links are 60%, 50%, 70%, and 40% respectively, then... , Generally, the lower the network load entropy value, the more balanced the link load.
[0055] S101. Based on the computing load in the computing node status, obtain the computing load entropy, which is used to measure the balance of computing nodes.
[0056] In one implementation, the expression for calculating the computing power load entropy can be:
[0057] in, It can refer to the entropy of computing power load. The percentage of FLOPS utilization of computing node v can be used to indicate the utilization rate of computing power node v. This can refer to a set of computing power nodes. For example, if the utilization rates of four computing power nodes are 70%, 60%, 50%, and 80%, then... , Generally, the lower the value of the computing load entropy, the more balanced the load of the computing nodes.
[0058] S102. Based on the proportion of service types in the service characteristics, the service fusion entropy is obtained. The service fusion entropy is used to measure the service compatibility within the same optical path data unit.
[0059] In one implementation, the expression for calculating the service load entropy can be:
[0060] in, It can refer to the entropy of business integration. This can refer to the proportion of service type τ within container k of ODUk. For example, if synchronous services currently account for 60% and independent services account for 40% within ODU0, then... , Generally, the lower the business load entropy value, the lower the business compatibility.
[0061] S103. Based on the network load entropy, computing power load entropy, and service fusion entropy, obtain the current coupling entropy.
[0062] In one implementation, the expression for calculating the current coupling entropy can be:
[0063] in, It can represent the current coupling entropy, and can represent the degree of disorder in the coupling of "network-computing power-service". It can represent the weighting coefficient of network load entropy. It can represent the weighting coefficient of computing load entropy. This can represent the weighting coefficients of the business integration entropy. Among them, and Optionally, For example, .
[0064] S104. Obtain the scheduling strategy based on the current coupling entropy and local strategy model.
[0065] In one optional embodiment, obtaining the scheduling strategy based on the current coupling entropy and the local policy model may include: constructing the current computing network state based on the current coupling entropy, bandwidth requirements in the network link state, and computing power requirements in the computing node state; calculating the similarity between the current computing network state and historical computing network states; if the similarity is greater than a state threshold, obtaining the scheduling strategy corresponding to the current computing network state based on the historical scheduling strategy corresponding to the historical computing network state; otherwise, inputting the current computing network state into the local policy model to generate the scheduling strategy for the current computing network state with the goal of maximizing the scheduling reward constructed based on the change in coupling entropy, the comprehensive utilization rate of computing network resources, and transmission latency; the computing network resource utilization rate is obtained based on the utilization rate of payload block resources, the utilization rate of ODUk resources, and the utilization rate of computing nodes.
[0066] In this embodiment, the local policy network can adopt a migration-generated scheduling strategy, with coupling entropy as the core, to achieve low-complexity and fast-convergence scheduling decisions.
[0067] This local policy network may include experience caching units, entropy-like migration units, entropy-differential generation units, and reward calculation units.
[0068] Historical scheduling experience can be stored in the experience cache unit. ,in This can represent the network state at time t. The network state at time t+1 can be represented. In one implementation, the network state may include coupling entropy, bandwidth requirements, and computing power requirements, for example... ,in, It can represent the coupling entropy at time t. This can represent the bandwidth requirement at time t. It can represent the computing power requirement at time t. The state of the computing network can also include other information, and is not limited to this. The scheduling action can represent a selection of payload blocks, optical path data units, routing information, and computing nodes. The final selected scheduling action serves as the scheduling strategy. This can represent a reward, which is positively correlated with the decrease in coupling entropy and resource utilization, and negatively correlated with latency.
[0069] Entropy-like transfer units can be used to identify historical network states that are highly similar to the current network state, and transfer the corresponding historical scheduling strategies as prior knowledge to the current scenario, thereby accelerating the generation and optimization of the current scheduling strategy.
[0070] Entropy variation generation units can be used to calculate scheduling strategies for current network states that do not have high similarity to historical network states.
[0071] Whether entropy-like migration or entropy-like generation can be determined by calculating the current network state. and historical network status The similarity is determined, and the expression for this similarity can be:
[0072] in, It can represent state similarity. It can represent the coupling entropy of the current computing network state. The coupling entropy can represent the historical state of the computing network. It can represent the 2-norm. This can represent the maximum coupling entropy, for example, It can be 2, used to normalize the similarity to the [0,1] interval.
[0073] If the state similarity Sim is greater than the preset state threshold, entropy-like migration can be selected, that is, the historical scheduling strategy of directly migrating the historical network state can be adopted. The initial scheduling direction is used to reduce redundant exploration; the scheduling strategy for the current computing network state is obtained based on this historical scheduling strategy.
[0074] If the state similarity Sim is less than a preset state threshold, entropy variation generation can be selected. In one implementation, the scheduling strategy for the current computing network state can be generated based on the relative entropy (Kullback–Leibler, KL) divergence, and its expression can be:
[0075]
[0076]
[0077]
[0078] in, This can be used to distribute pre-trained baseline scheduling policies, avoiding instability caused by over-exploration of the generated scheduling policies and ensuring that the generated scheduling policies balance adaptability and stability. This can be the distribution of scheduling strategies for the current computing network state. Can be The scheduling action, Can be The scheduling action, This can represent the scheduling reward at time t. The reward-weighted coefficient can represent the difference in coupling entropy. It can represent the coupling entropy before the execution of a scheduling action. It can represent the coupling entropy after executing a scheduling action. The reward-weighted coefficient can represent the comprehensive utilization rate of computing network resources. It can represent the comprehensive utilization rate of computing network resources. A reward-weighted coefficient that can represent the difference in latency. This can represent the service transmission delay of the current scheduling path at time t. It can represent the minimum achievable transmission delay in a network. This can represent the maximum possible transmission delay in the network. This can represent the weighting coefficient of the load block resource utilization rate. This can represent the utilization rate of the load block resources. This can represent the weighting coefficient of ODUk resource utilization. It can represent the utilization rate of ODUk resources. It can represent the weighting coefficient of computing node utilization. It can represent the utilization rate of computing nodes.
[0079] Further optionally, obtaining the scheduling strategy corresponding to the current computing network state based on the historical scheduling strategy corresponding to the historical computing network state may include: adjusting the historical scheduling strategy according to the service characteristics to obtain the scheduling strategy corresponding to the current computing network state.
[0080] For example, extracting from the experience cache =1.15 historical state Calculate similarity It is determined to be an "entropy-like state"; historical scheduling strategy is migrated. This historical scheduling strategy To select 3 70Mbps PBs, map OPU to OPU1, select ODUflex for ODUk, route (0,0)→(1,0)→(2,0)→(2,1)→(2,2), and select v3 as the computing node; combined with the current service G1 priority (high), fine-tune the PB time slots to continuous time slots 1-21 to ensure low latency.
[0081] In an optional embodiment, the frame information is obtained by processing service data, local state information, and scheduling strategy, and may include: determining the target optical path load unit based on the service characteristics in the local state information; synchronous services in the same service group are located in the same target optical path load unit, and the synchronous services are determined by service characteristics; mapping service data and local state information to the target load block; mapping the target load block to the target optical path load unit; the number of load blocks in the target optical path load unit does not exceed a number threshold; mapping the target optical path load units of the same synchronous service to the same target optical path data unit; the target load block occupies a continuous time slot position in the target optical path data unit; and mapping the target optical path data unit to the target optical path transmission unit to obtain frame information.
[0082] like Figure 3 The multi-layered mapping relationship shown on the right, in the PB-OPU mapping layer: can be adjusted according to bandwidth requirements. Selecting the size of the fine-grained unit of the optical network (OPU) determines the target payload block. For example, the bandwidth size of the fine-grained unit is incremented in 10Mbps steps. A payload block (PB) can contain seven 10Mbps physical time slots, meaning the full capacity of a single PB is 70Mbps. For example, when the bandwidth requirement is no greater than 70Mbps, consecutive physical time slots within a single target payload block can be selected. For instance, for a 30Mbps service requirement, three consecutive physical time slots within a single PB can be allocated. When the bandwidth requirement exceeds 70Mbps, multiple target payload blocks are selected. For instance, for a 100Mbps service requirement, all seven physical time slots of the first target payload block and three consecutive physical time slots within the second target payload block are allocated. It is ensured that a single OPU aggregates a maximum of seven PBs, and synchronous services within the same service group are preferentially mapped to the same OPU. At the OPU-ODUk mapping layer, coupling entropy can be used to determine the target payload block size. Alternatively, service convergence entropy can be used to map OPUs to ODUk containers, for example, selecting from ODU0, ODU1, or ODUflex. The constraints are that OPUs belonging to synchronous services are mapped to the same ODUk, and the PB time slots within the ODUk are continuous, avoiding "time slot holes." At the ODUk-OTU mapping layer, ODUk can be mapped to the wavelength of the target route to obtain frame information. This target route can be determined based on network load entropy. You get the choice.
[0083] In this application embodiment, a schematic diagram of frame information format is provided, such as... Figure 3As shown on the left, FA in the diagram represents signal frame alignment overhead, OTUK represents optical channel transmission unit overhead, ODUK represents optical channel data unit overhead, OPUK represents optical channel payload unit overhead, PB represents payload block, and n represents the identification information of the payload block, indicating different payload blocks. For example, PB#1 can represent the payload block identified as 1. The definitions of PB#8, PB#n-1, and PB#n in the diagram can be found in the above description and will not be elaborated further here. In this frame information format, there are n payload blocks. Each payload block can encapsulate service data and local state information. In one implementation, the payload block can encapsulate a fine-grained flexible optical service unit payload block frame (OSUflex PB frame). This OSUflex PB frame may include the following information: Version Identifier (VER), Computing Power Optical Network Transmission File Protocol (CPO-TPN), Reserved Fields (RES), Channel Monitoring-Serial Connection Monitoring (PM-TCM), Path Tracing Identifier (TT1), 32-Frame Indicator (M32), Automatic Protection Switching (APS), Delay Measurement (DM), Integrated Monitoring (MON), Continuity Check (CV), Sequence Number (SQ), Mapping Overhead, Cyclic Redundancy Check (CRCB), and Payload. This can represent frame information. This diagram is only an example of a frame format and is not limited to it.
[0084] Furthermore, during the Acknowledgement (ACK) transmission phase, the optical path is locked to prevent other services from occupying it, ensuring the timing synchronization of computing power service transmission.
[0085] For example, in the PB-OPU mapping layer, three 70Mbps PBs (total 210Mbps) are mapped to OPU1, occupying time slots 1-21 of OPU1; in the OPU-ODUk mapping layer, OPU1 is mapped to ODUflex (capacity 250Mbps), ensuring that ODUflex only contains the synchronization service belonging to G1; in the ODUk-OTU mapping layer, ODUflex is mapped to wavelength 3, with the route (0,0)→(1,0)→(2,0)→(2,1)→(2,2); ACK information is sent to lock the optical path and prohibit other services from accessing.
[0086] In one optional embodiment, after acquiring frame information, the source node modulates it to a specific wavelength via electro-optical conversion (E / O conversion) and transmits it to the next node; the receiving node, after acquiring the wavelength signal, restores it to the original frame information using optical-electrical conversion (O / E conversion). With successful reception and parsing of the confirmation information, an optical path unlocking command can be automatically triggered, releasing the physical wavelength resources previously monopolized to ensure timing synchronization. This allows the optical path to return to the dynamic resource scheduling pool, achieving immediate recovery of network carrying efficiency and resource recycling while ensuring strong timing consistency of computing tasks. Furthermore, the scheduling strategy can be stored as scheduling experience in an experience cache, and based on this, the parameters of the source node's local strategy model can be updated.
[0087] In an optional embodiment, the method may include: dividing the original service data into multiple time-correlated synchronous service data based on service characteristics. For the service data of multiple synchronous services, a local policy model can be used to calculate their respective scheduling policies, thereby obtaining corresponding frame information, which are then scheduled to the same or different destination nodes for computation and processing. An ACK information locking optical path mechanism can be used to maintain the time synchronization between the nodes.
[0088] The following example illustrates the scheduling of AI inference services under a computing power optical network topology: 1. Topology and Service Parameters Topology: Mesh-structured optical computing power network with fine-grained unit optical computing power, computing power nodes ( There are 16 nodes in total (coordinates (x,y), x,y=0~3), each node has a computing power capacity of 1000FLOPS; there are 24 optical links (E), each link has 80 wavelengths (100Gbps / wavelength), and the computing power optical network fine-grained unit -PB is 70Mbps (7×10Mbps time slots).
[0089] Business: AI inference business G1 (synchronous, high priority), source node (0,0), destination node (2,2), bandwidth requirement 200Mbps, computing power requirement 500FLOPS, latency tolerance ≤5ms.
[0090] 2. Scheduling and execution process Local status information collected: 1. Network links: Utilization of link (0,0)-(1,0) is 50%, (1,0)-(2,0) is 45%, (2,0)-(2,1) is 60%, and (2,1)-(2,2) is 40%; 2. Computing nodes: Node (0,0) has a FLOPS utilization of 60% (600 FLOPS / 1000 FLOPS), and node (2,2) has a utilization of 50% (500 FLOPS / 1000 FLOPS); 3. Service G1 parameters: =200Mbps =500 FLOPS =Synchronous.
[0091] Coupling entropy calculation: 1. Based on the utilization rates of the four candidate links, the following calculations were performed: 2. Based on the utilization rates of (0,0), (2,2) and surrounding nodes, the following calculations were performed: 3. There are currently no synchronization services within ODUflex. 4. =0.4×1.36+0.3×1.32+0.3×0=0.94.
[0092] Migration generation scheduling: 1. Matching in the experience cache =0.92 historical network state (Sim=0.99>0.7), migration actions: PB selects 3 70Mbps (210Mbps), OPU is mapped to OPU2, ODUk selects ODUflex, route (0,0)→(1,0)→(2,0)→(2,1)→(2,2); 2. Fine-tuning: Since G1 bandwidth is 200Mbps≤210Mbps, retain the PB configuration; prioritize the (2,2) node (50% utilization, meeting the 500FLOPS requirement).
[0093] Multi-layer mapping execution: 1. PB-OPU: 3 PBs (time slots 1-7, 8-14, 15-21) are mapped to OPU2; 2. OPU-ODUk: OPU2 is mapped to ODUflex (capacity 250Mbps), locking only synchronous services within ODUflex; 3. ODUk-OTU: ODUflex is mapped to wavelength 5, opening the link (0,0)-(1,0)-(2,0)-(2,1)-(2,2), locking the optical path; 4. ACK transmission: ACK is returned from (2,2) to (0,0), keeping the optical path locked.
[0094] Computing power service transmission: 1. AI inference data is transmitted along wavelength 5 after E / O conversion, with an end-to-end latency of 3.2ms (meeting the ≤5ms requirement); 2. After transmission is complete, the optical path is unlocked, and the experience ( =0.94, =0.85, =0.88) is stored in the cache, and the parameters of the local policy model are updated.
[0095] 3. Implementation Results In terms of resource utilization, the PB utilization rate is 75% (210Mbps / 280Mbps OPU capacity), and the utilization rate of computing nodes (2,2) is 100% (500FLOPS+500FLOPS). Regarding latency performance, the end-to-end latency is 3.2ms, a 45% reduction compared to traditional ODU0 scheduling (5.8ms); and in terms of convergence speed, the scheduling decision time is 12ms, an 89% reduction compared to dynamic multi-DQN (110ms).
[0096] like Figure 4 As shown, in another optional embodiment, a data processing method is provided, applied to an intermediate node, and may include: S1b: Receive local state information perceived by the intermediate node, frame information sent by the previous node, and scheduling policy; the previous node is either the source node or the intermediate node; the frame information includes service data and local state information of the upstream node; the local state information includes service characteristics, network link status, and computing power node status; the scheduling policy of the previous node includes target payload block, target optical path data unit, target routing information, and target computing power node, which is obtained by dynamically updating the scheduling policy of the source node along the path during transmission based on the local state information of the upstream node and the corresponding local policy model.
[0097] S2b: Decapsulate the frame information according to the scheduling strategy of the previous node to extract the business data and the local state information of the upstream node.
[0098] S3b: Calculate a new scheduling strategy based on the local state information of the upstream node, the local state information perceived by the intermediate node, and the received local strategy model.
[0099] S4b: Based on business data, local state information of intermediate nodes, and new scheduling strategies, new frame information is obtained.
[0100] In step S3b, when generating a new scheduling strategy, the path planning of downstream links can be considered, while the path of upstream links remains unchanged. The new scheduling strategy may be the same as or different from the original scheduling strategy. In step S4b, when generating new frame information, the local state information of intermediate nodes can overwrite the local state information of upstream nodes, or the local state information of intermediate nodes can be added to the local state information of upstream nodes as the local state information of the next node's upstream nodes. This can be understood as the local state information of upstream nodes being either the local state information of the previous node or the local state information of at least some upstream nodes. The implementation method for obtaining new frame information can be referred to the above description, and will not be elaborated further here.
[0101] In this embodiment, after receiving the frame information sent by the previous node, the intermediate node does not need to send the received frame information and its own state information to the controller to wait for the controller to make a scheduling decision. Instead, the node processes the local state information in the frame information and the local state information it perceives based on its latest local policy model to achieve frame-level response.
[0102] In one optional embodiment, a new scheduling strategy is calculated based on the local state information of the upstream node, the local state information perceived by the intermediate node, and the received local policy model. This may include: constructing the current joint along-path computing network state based on the local state information of the upstream node and the local state information perceived by the intermediate node; calculating the similarity between the current joint along-path computing network state and the historical joint along-path computing network state; and obtaining the new scheduling strategy based on the similarity. The implementation methods for constructing the computing network state, calculating the similarity, and obtaining the scheduling strategy based on the similarity can be referred to the foregoing description, and will not be elaborated further here.
[0103] like Figure 5 As shown, in another optional embodiment, a data processing method is provided, applied to a destination node, and may include: S1c: Receive the frame information and corresponding scheduling policy from the previous node; the previous node is either the source node or an intermediate node; the frame information includes service data and the local state information of the upstream node, and the scheduling policy includes the target payload block, the target optical path data unit, the target routing information, and the target computing power node. The scheduling policy of the previous node is obtained by dynamically updating the scheduling policy of the source node along the path based on the local state information of the upstream node and the corresponding local policy model during the transmission process.
[0104] S2c decapsulates the frame information according to the scheduling strategy of the previous node, and extracts the business data and the local state information of the upstream node.
[0105] S3c sends business data to the target computing node for processing to obtain the computing power execution result.
[0106] In an optional embodiment, the method further includes: receiving the computing node status reported by the target computing node; obtaining the local state information of the target node based on the computing node status; and reporting the local state information of the target node to the controller for updating the global policy model, wherein the updated global policy model is used to update the local policy model.
[0107] In an optional embodiment, the method further includes: monitoring the execution status of the target computing node on the business data; if the business data has not been processed, the target node is made into a new source node, and the scheduling policy and frame information are recalculated and generated based on its current perceived local state information and local policy model, so as to unload the remaining business data to the subsequent computing node for execution.
[0108] The specific implementation of the embodiments of this application can be referred to the description of the foregoing embodiments, and will not be repeated here.
[0109] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.
[0110] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0111] Based on the same inventive concept, corresponding to any of the methods in the above embodiments, this application also provides a data processing apparatus, which may include: The receiving module is used to receive local state information and local policy models.
[0112] The calculation module is used to calculate a scheduling strategy based on the local state information and the local policy model. The local state information includes service characteristics, network link status, and computing node status. The scheduling strategy includes target payload blocks, target optical path data units, target routing information, and target computing nodes. The local policy model is obtained by the controller based on the global policy model. Frame information is obtained by processing the service data, the local state information, and the scheduling strategy. During transmission, the scheduling strategy is updated by the receiving node based on its own perceived local state information and the corresponding local policy model.
[0113] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.
[0114] The apparatus of the above embodiments is used to implement the corresponding data processing method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0115] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the data processing method described in any of the above embodiments.
[0116] Figure 6 This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0117] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0118] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0119] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0120] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0121] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0122] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0123] The electronic devices described above are used to implement the corresponding data processing methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0124] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the data processing method as described in any of the above embodiments.
[0125] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0126] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the data processing method of any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0127] It should be noted that the embodiments of this application can also be further described in the following ways: A data processing method, applied to a source node, may include: calculating a scheduling strategy based on local state information perceived by the source node and a received local policy model; the local state information includes service characteristics, network link status, and computing node status; the scheduling strategy includes target payload blocks, target optical path data units, target routing information, and target computing nodes; the local policy model is obtained by the controller based on a global policy model; frame information is obtained by processing service data, local state information, and the scheduling strategy; and the scheduling strategy is updated along the path by the receiving node during transmission based on the perceived local state information and the corresponding local policy model.
[0128] Optionally, a scheduling strategy is calculated based on the local state information perceived by the source node and the received local policy model, including: obtaining network load entropy based on link utilization in the network link state, which is used to measure link load balance; obtaining computing load entropy based on computing load in the computing node state, which is used to measure computing node balance; obtaining service fusion entropy based on the proportion of service types in service characteristics, which is used to measure service compatibility within the same optical path data unit; obtaining the current coupling entropy based on network load entropy, computing load entropy, and service fusion entropy; and obtaining the scheduling strategy based on the current coupling entropy and the local policy model.
[0129] Optionally, a scheduling strategy is obtained based on the current coupling entropy and the local policy model, including: constructing the current computing network state based on the current coupling entropy, bandwidth requirements in the network link state, and computing power requirements in the computing node state; calculating the similarity between the current computing network state and historical computing network states; if the similarity is greater than a state threshold, obtaining the scheduling strategy corresponding to the current computing network state based on the historical scheduling strategy corresponding to the historical computing network state; otherwise, inputting the current computing network state into the local policy model to generate a scheduling strategy for the current computing network state with the goal of maximizing the scheduling reward constructed based on the change in coupling entropy, the comprehensive utilization rate of computing network resources, and transmission latency; the computing network resource utilization rate is obtained based on the utilization rate of payload block resources, the utilization rate of ODUk resources, and the utilization rate of computing nodes.
[0130] Optionally, frame information is obtained by processing service data, local state information, and scheduling strategies, including: determining the target optical path load unit based on service characteristics in the local state information; synchronous services in the same service group are located in the same target optical path load unit, and synchronous services are determined by service characteristics; mapping service data and local state information to the target load block; mapping the target load block to the target optical path load unit; ensuring that the number of load blocks in the target optical path load unit does not exceed a certain threshold; mapping the target optical path load units of the same synchronous service to the same target optical path data unit; ensuring that the target load block occupies a continuous time slot position in the target optical path data unit; and mapping the target optical path data unit to the target optical path transmission unit to obtain frame information.
[0131] A data processing method, applied to an intermediate node, may include: receiving local state information perceived by the intermediate node, frame information sent by the previous node, and a scheduling policy; the previous node is either a source node or an intermediate node; the frame information includes service data and local state information of the upstream node; the local state information includes service characteristics, network link status, and computing node status; the scheduling policy of the previous node includes target payload blocks, target optical path data units, target routing information, and target computing nodes, which is obtained by dynamically updating the scheduling policy of the source node along the path during transmission based on the local state information of the upstream node and the corresponding local policy model; decapsulating the frame information according to the scheduling policy of the previous node to extract the service data and the local state information of the upstream node; calculating a new scheduling policy based on the local state information of the upstream node, the local state information perceived by the intermediate node, and the received local policy model; and processing the service data, the local state information of the intermediate node, and the new scheduling policy to obtain new frame information.
[0132] A data processing method, applied to a destination node, may include: receiving frame information and a corresponding scheduling policy from a previous node; the previous node may be a source node or an intermediate node; the frame information includes service data and local state information of the upstream node, and the scheduling policy includes a target payload block, a target optical path data unit, target routing information, and a target computing power node; the scheduling policy of the previous node is obtained by dynamically updating the scheduling policy of the source node along the path during transmission based on the local state information of the upstream node and the corresponding local policy model; decapsulating the frame information according to the scheduling policy of the previous node to extract the service data; and sending the service data to the target computing power node for processing to obtain the computing power execution result.
[0133] Optionally, the method further includes: receiving the computing node status reported by the target computing node; obtaining the local state information of the target node based on the computing node status; and reporting the local state information of the target node to the controller for updating the global policy model, wherein the updated global policy model is used to update the local policy model.
[0134] Optionally, the method further includes: monitoring the execution status of the target computing node on the business data; if the business data has not been processed, the target node is made into a new source node, and the scheduling policy and frame information are recalculated and generated based on its current perceived local state information and local policy model, so as to unload the remaining business data to the subsequent computing node for execution.
[0135] A data processing apparatus may include: a receiving module for receiving local state information and a local policy model; and a calculation module for calculating a scheduling policy based on the local state information and the local policy model. The local state information includes service characteristics, network link status, and computing node status. The scheduling policy includes target payload blocks, target optical path data units, target routing information, and target computing nodes. The local policy model is obtained by the controller based on a global policy model. Frame information is obtained by processing the service data, local state information, and scheduling policy. During transmission, the scheduling policy is updated by the receiving node based on its own perceived local state information and the corresponding local policy model.
[0136] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the aforementioned data processing method.
[0137] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.
[0138] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0139] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0140] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
Claims
1. A data processing method, characterized in that, Applied to the source node, including: The scheduling strategy is calculated based on the local state information perceived by the source node and the received local strategy model; the local state information includes service characteristics, network link status, and computing node status; the scheduling strategy includes target payload block, target optical path data unit, target routing information, and target computing node; the local strategy model is obtained by the controller based on the global strategy model; Frame information is obtained by processing business data, the local state information, and the scheduling strategy. The scheduling strategy is updated along the path by the receiving node during transmission based on the perceived local state information and the corresponding local strategy model.
2. The method according to claim 1, characterized in that, The scheduling policy is calculated based on the local state information perceived by the source node and the received local policy model, including: Based on the link utilization in the network link status, the network load entropy is obtained, which is used to measure the link load balance. The computing load entropy is obtained based on the computing load degree in the computing node status. The computing load entropy is used to measure the balance of computing nodes. The service fusion entropy is obtained based on the proportion of service types in the service characteristics. The service fusion entropy is used to measure the service compatibility within the same optical path data unit. The current coupling entropy is obtained based on the network load entropy, the computing power load entropy, and the service fusion entropy; The scheduling strategy is obtained based on the current coupling entropy and the local strategy model.
3. The method according to claim 2, characterized in that, The scheduling policy is obtained based on the current coupling entropy and the local policy model, including: The current computing network state is constructed based on the current coupling entropy, the bandwidth requirement in the network link state, and the computing power requirement in the computing power node state; Calculate the similarity based on the current computing network state and the historical computing network state; If the similarity is greater than the state threshold, then the scheduling strategy corresponding to the current computing network state is obtained based on the historical scheduling strategy corresponding to the historical computing network state. Otherwise, the current computing network state is input into the local policy model to generate a scheduling policy for the current computing network state with the goal of maximizing the scheduling reward constructed based on the change in coupling entropy, the comprehensive utilization rate of computing network resources, and the transmission delay; the computing network resource utilization rate is obtained based on the utilization rate of payload block resources, the utilization rate of optical path data unit resources, and the utilization rate of computing power nodes.
4. The method according to claim 1, characterized in that, Frame information is obtained by processing business data, the local state information, and the scheduling strategy, including: The target optical path load unit is determined based on the service characteristics in the local state information; synchronous services in the same service group are located in the same target optical path load unit, and the synchronous services are determined by the service characteristics. Map the business data and the local state information into the target payload block; The target payload block is mapped into the target optical path payload unit; the number of payload blocks in the target optical path payload unit does not exceed a quantity threshold. The target optical path payload unit of the same synchronous service is mapped to the same target optical path data unit; the target payload block occupies a continuous time slot position in the target optical path data unit; The target optical path data unit is mapped to the target optical path transmission unit to obtain the frame information.
5. A data processing method, characterized in that, Applied to intermediate nodes, including: The system receives local state information perceived by the intermediate node, frame information sent by the previous node, and scheduling policy; the previous node is either a source node or an intermediate node; the frame information includes service data and local state information of the upstream node; the local state information includes service characteristics, network link status, and computing power node status; the scheduling policy of the previous node includes target payload block, target optical path data unit, target routing information, and target computing power node, and is obtained by dynamically updating the scheduling policy of the source node along the path during transmission based on the local state information of the upstream node and the corresponding local policy model; The frame information is decapsulated according to the scheduling strategy of the upstream node to extract the service data and the local state information of the upstream node; A new scheduling strategy is calculated based on the local state information of the upstream node, the local state information perceived by the intermediate node, and the received local strategy model. New frame information is obtained by processing the business data, the local state information of the intermediate node, and the new scheduling strategy.
6. A data processing method, characterized in that, Applied to the destination node, including: The system receives frame information and corresponding scheduling policy from the previous node; the previous node is either a source node or an intermediate node; the frame information includes service data and local state information of the upstream node; the scheduling policy includes target payload block, target optical path data unit, target routing information and target computing power node; the scheduling policy of the previous node is obtained by dynamically updating the scheduling policy of the source node along the path based on the local state information of the upstream node and the corresponding local policy model during transmission. The frame information is decapsulated according to the scheduling strategy of the previous node to extract the service data; The business data is sent to the target computing node for processing to obtain the computing power execution result.
7. The method according to claim 6, characterized in that, The method further includes: Receive the computing node status reported by the target computing node; Based on the state of the computing power node, the local state information of the target node is obtained; The local state information of the target node is reported to the controller to update the global policy model, and the updated global policy model is used to update the local policy model.
8. The method according to claim 6, characterized in that, The method further includes: Monitor the execution status of the target computing power node on the business data; If the business data is not fully processed, the destination node is made a new source node. Based on its current perceived local state information and local strategy model, the scheduling strategy and frame information are recalculated and generated to unload the remaining business data to subsequent computing power nodes for execution.
9. A data processing apparatus, characterized in that, include: The receiving module is used to receive local state information and local policy models; The calculation module is used to calculate a scheduling strategy based on the local state information and the local policy model. The local state information includes service characteristics, network link status, and computing node status. The scheduling strategy includes target payload blocks, target optical path data units, target routing information, and target computing nodes. The local policy model is obtained by the controller based on the global policy model. Frame information is obtained by processing the service data, the local state information, and the scheduling strategy. During transmission, the scheduling strategy is updated by the receiving node based on its own perceived local state information and the corresponding local policy model.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as claimed in any one of claims 1 to 8.