Link reconstruction method and device and cluster

The controller determines the port connection relationship based on the computing power parameters of the switch unit, sends link allocation information to the optical space switching device, realizes network topology reconstruction configuration, solves the problems of increased cost of optical space switching devices and data packet loss in the prior art, and realizes lower cost and more efficient network topology management.

CN120017607APending Publication Date: 2025-05-16CHENGDU HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311515852.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art realizes dynamic adjustment of network topology by monitoring network traffic distribution during task execution, resulting in increased cost of optical space switching equipment and data packet loss problems.

Method used

The computing power parameters of the switch unit are obtained through the controller, the port connection relationship between the switch units is determined, and the link allocation information is sent to the optical space switching device, so as to realize network topology reconstruction configuration, avoiding the complexity of monitoring network traffic and additional storage requirements.

Benefits of technology

It reduces the computing complexity and hardware cost of optical space switching devices, and avoids data packet loss and network resource waste caused by link switching during task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017607A_ABST
    Figure CN120017607A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a link reconstruction method, a link reconstruction device and a cluster. A controller obtains resource information of each switch group for indicating service execution in an all-optical network and computing power parameters of each switch group; when the controller receives the resource information of the service, determining link allocation information used for indicating a port connection relationship between the switching units executing the service according to the computing power parameter of each switching unit, and sending the link allocation information to the optical space switching equipment in the all-optical network. The port connection among the switch sets is allocated according to the resource information of the service, network topology reconstruction configuration is realized, the distribution condition of network flow does not need to be monitored, the cost required by the switch and the controller can be reduced, and meanwhile, the reconstruction configuration of links in the network is performed before the service is executed, so that the network topology reconstruction configuration efficiency is improved. The problem of data packet loss caused by link switching in the task execution process and the problem of network resource waste caused by waiting for emptying of the data flow on the link are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technology, and more specifically, to a link reconstruction method, device, and cluster. Background Art

[0002] With the development trend of Moore's Law, the development of network performance has gradually failed to keep up with the development of computing performance, becoming a performance bottleneck. In order to effectively improve network performance, commonly used methods include expanding switch capacity, optimizing routing algorithms and congestion control algorithms, etc., by adjusting the distribution of load in the network, load balancing is achieved as much as possible to improve network performance. In addition, starting from the network topology itself, micro-electro-mechanical-system-based optical switches (MEMS-based optical switch, MBOS) or optical cross-connect (OXC) and other optical space switching devices can be introduced to increase network capacity while dynamically adjusting the network topology to achieve flexible scheduling of network bandwidth resources.

[0003] At present, the existing related technologies mainly involve the optical space switching device monitoring the distribution of network traffic during task execution, establishing a new link connection relationship based on the distribution of network traffic, and performing link switching and routing configuration according to the new link connection relationship to achieve dynamic adjustment of the network topology. Due to the high complexity and operation of network traffic monitoring, and the different iteration cycles of different tasks, the monitoring and topology switching time is difficult to determine, and the need to store historical traffic information additionally requires adding additional storage devices to the optical space switching device used to control link switching, or selecting an optical space switching device with a larger storage space, which will increase the hardware cost of the optical space switching device, and there is a risk of data packet loss when performing link switching during task execution. Summary of the invention

[0004] The present application provides a link reconstruction method, device and cluster to solve the problem of increased cost of optical space switching equipment and data packet loss caused by the prior art of dynamically adjusting the network topology structure by monitoring the distribution of network traffic during task execution.

[0005] In a first aspect, a link reconstruction method is proposed for use in an all-optical network, the method being executed by a controller or a functional module or chip in the controller, the method comprising: obtaining resource information indicating each switch group executing a service in the all-optical network and a computing power parameter of each switch group, determining link allocation information indicating a port connection relationship between switch groups executing services according to the computing power parameters of each switch group, and sending the link allocation information to an optical space switching device in the all-optical network. Optionally, the computing power parameter can be used to indicate the computing power of the switch group when executing a service.

[0006] Based on the method described in the first aspect, when the controller receives the resource information of the service, it allocates the port connection between the switch groups according to the resource information of the service to realize the network topology reconstruction configuration. The optical space switching device does not need to monitor the distribution of network traffic. Instead, the controller determines the link allocation information and sends it to the optical space switching device, and the optical space switching device performs the link reconstruction operation based on the link allocation information. This can reduce the computational complexity of the optical space switching device, and further reduce the hardware cost of the optical space switching device. At the same time, the reconstruction configuration of the links in the network is performed before the service is executed, so as to avoid the data packet loss caused by link switching during the task execution, and the network resource waste caused by waiting for the data flow on the link to be drained.

[0007] In one possible implementation, the port connection relationship is used to indicate or include the number of connection ports and connection port identifiers between each switch group. The port connection relationship can be established based on the number of link allocations between each switch group and the port information of each switch group; the number of connection ports between each switch group is the same as the number of link allocations between each switch group; the number of link allocations between each switch group is allocated according to the computing power parameter ratio of each switch group.

[0008] Based on this possible implementation method, the number of link allocations between each switch group is allocated according to the computing power parameter ratio of each switch group. There is no need to monitor the traffic in the network, thus avoiding the traffic monitoring cost. At the same time, according to the computing power parameter ratio of different switch groups, the network link utilization rate can be improved while achieving balanced link allocation.

[0009] In one possible implementation, the computing power parameter of the switch group may include the number of computing nodes occupied by the switch group when executing the business. Accordingly, the above-mentioned example parameter ratio can be understood as the ratio of the number of computing nodes occupied when executing the business, and the number of link allocations between each switch group is allocated according to the ratio of the number of computing nodes occupied by each switch group when executing the business.

[0010] Based on this possible implementation method, the number of link allocations between each switch group is allocated according to the number of computing nodes occupied by each switch group when executing the business. There is no need to monitor the traffic in the network, thus avoiding the cost of traffic monitoring. At the same time, according to the distribution ratio of the number of computing nodes occupied by different switch groups when performing the business, it is possible to improve the utilization rate of network links and achieve balanced link allocation.

[0011] In a possible implementation, allocating the number of link allocations between the switch groups according to the ratio of the number of computing nodes occupied by each switch group when executing the service is specifically implemented as follows: selecting the first switch group from the switch groups executing the service in order from low to high in the number of computing nodes occupied by each switch group when executing the service, determining the link allocation ratio between the switch groups according to the ratio between the number of computing nodes occupied by the first switch group when executing the service and the number of computing nodes occupied by the remaining switch groups when executing the service, and obtaining the link allocation number between the switch groups according to the link allocation ratio between the switch groups and the maximum link allocation number of each switch group; the maximum link allocation number of the switch group is equal to the number of computing nodes occupied by the switch group when executing the service.

[0012] Based on this possible implementation, link allocation is started from the switch group that occupies the lowest number of computing nodes when executing the service, so as to achieve balanced link allocation among the switch groups.

[0013] In a possible implementation, obtaining the link allocation quantity between each switch group according to the link allocation ratio between each switch group and the maximum link allocation quantity of each switch group is specifically implemented as follows: according to the link allocation ratio between the switch group and each second switch group, the maximum link allocation quantity of the switch group is allocated to each second switch group to obtain the link allocation quantity between the switch group and each second switch group; the second switch group is a switch group other than the switch group in the switch group executing the service; the sum of the link allocation quantities between the switch group and each second switch group is less than or equal to the maximum link allocation quantity of the switch group.

[0014] Based on this possible implementation method, the maximum link allocation number is used to ensure that the link data that can be allocated to each switch group is less than or equal to the number of computing nodes occupied by the switch group when executing the business, and link allocation starts from the switch group with the lowest number of computing nodes occupied when executing the business, so as to achieve balanced link allocation between the switch groups.

[0015] In a possible implementation, allocating the maximum link allocation quantity of the switch group to each second switch group to obtain the link allocation quantity between the switch group and each second switch group is specifically implemented as follows: allocating the maximum link allocation quantity of the switch group to each second switch group according to the link allocation ratio between the switch group and each second switch group, and determining the initial link allocation quantity between the switch group and each second switch group; if the sum of the initial link allocation quantities between the switch group and each second switch group is equal to the maximum link allocation quantity of the switch group, then determining the initial link allocation quantity between the switch group and each second switch group as the link allocation quantity between the switch group and each second switch group; if the sum of the initial link allocation quantities between the switch group and each second switch group is less than the maximum link allocation quantity of the switch group, then adjusting the initial link allocation quantity between the switch group and the third switch group to obtain the link allocation quantity between the switch group and each second switch group; the third switch group is included in the second switch group whose link allocation quantity is less than the maximum link allocation quantity.

[0016] Based on this possible implementation method, when the controller allocates links according to the distribution ratio of the number of computing nodes occupied by the switch group when executing the business, when there are switch groups with remaining connection ports, the remaining connection ports are utilized by increasing the links between the switch groups with remaining connection ports to further improve the network link utilization.

[0017] In a possible implementation, the third switch group is selected from the second switch groups whose link allocation number is less than the maximum link allocation number according to the traffic characteristics between the switch group and each second switch group. Optionally, the traffic characteristics can be used to indicate the amount of data transmission between the switch groups when executing the service.

[0018] Based on this possible implementation method, when the controller allocates links according to the distribution ratio of the number of computing nodes occupied by the switch group when executing the business, it selects the switch group that uses the remaining connection ports according to the traffic characteristics between the switch groups, so as to further achieve traffic balance between the switch groups.

[0019] In a possible implementation, after sending link allocation information to an optical space switching device in an all-optical network, the method may further include: sending a routing configuration instruction to each switch group according to the link allocation information to instruct each switch group to perform routing configuration according to the link allocation information, so as to establish a routing path between each switch group.

[0020] In a possible implementation, the link allocation information is also used to indicate the connection port of each switch group in the switch group that executes the service; the configuration instruction may include an instance creation instruction and a port binding instruction; the instance creation instruction may be used to instruct the switch group to create a virtual routing instance corresponding to the service in the switch group, and the instance creation instruction is sent to the switch group that executes the service based on the switch group that executes the service indicated by the link allocation information; the port binding instruction may be used to instruct the switch group to bind the connection port of the switch group to the virtual routing instance corresponding to the service; the port binding instruction is sent to the switch group that executes the service based on the identification information of the connection port of each switch group after receiving the creation confirmation information returned by the switch group that executes the service based on the instance creation instruction.

[0021] Based on this possible implementation, a virtual routing instance corresponding to the service is created in the switch group, and routing isolation is implemented based on the virtual routing instance to avoid the phenomenon of link grabbing between services, which helps to alleviate network congestion.

[0022] In a possible implementation, the virtual routing instance includes a shortest virtual routing instance and a non-shortest virtual routing instance.

[0023] Based on this possible implementation, a non-shortest virtual route is deployed in the reconstructed network to reduce the congestion of the shortest path and solve the load balancing problem.

[0024] In a possible implementation, the service identifier of the service is associated with the link allocation information, and the method may further include: in response to a link release request for requesting the release of a link for transmitting data of a target service, obtaining a target service identifier of the target service carried in the link release request; determining the target link allocation information associated with the target service identifier, and sending a link deletion request to the optical space switching device based on the target link allocation information to instruct the optical space switching device to delete the port connection relationship indicated by the target link allocation information, thereby deleting the port connection relationship corresponding to the target link allocation information in the all-optical network; and sending a path deletion instruction to the target switch group of the target service to instruct the deletion of the target virtual routing instance corresponding to the target service in the target switch group.

[0025] Based on this possible implementation, the controller associates the service identifier with the link allocation information, multiple services do not share links, and the allocated link dedicated to a service can only be used to transmit the data of the service, so as to isolate network resources between different services and avoid the phenomenon of link grabbing between services, which helps to alleviate network congestion. And when it is necessary to release the link of the service, the port connection relationship indicated by the target link allocation information corresponding to the service identifier of the service can be deleted, without waiting for other services to complete or traffic migration, to avoid data packet loss.

[0026] In a possible implementation, obtaining service resource information may include: responding to a link application request, obtaining service resource information obtained by a scheduler based on service resources carried in the service application request and task scheduling of currently idle computing nodes in the all-optical network. Based on this possible implementation, service resource information may be obtained from the scheduler, reducing the complexity of the controller obtaining service resource information.

[0027] In the second aspect, a link reconstruction method is proposed, which is applied to an optical space switching device in an all-optical network or a functional module or chip in the optical space switching device, and the method includes: the optical space switching device receives link allocation information from a controller for indicating the port connection relationship between switch groups that execute services; the link allocation information is obtained by the controller through computing power parameters of each switch group; the computing power parameters of each switch group are obtained by the controller from resource information of each switch group that indicates the execution of services in the all-optical network and the computing power parameters of each switch group; the optical space switching device reconstructs and configures the optical path connection between the switch groups according to the link allocation information to form an optical path connection relationship between the optical space switching device and each switch group.

[0028] Based on the method described in the second aspect, the optical space switching device is connected to the controller, and based on the link allocation information sent by the controller, an optical path connection relationship between the optical space switching device and each switch group is established, and by responding to the link deletion request sent by the controller, the port connection relationship indicated by the target link allocation information carried in the link deletion request is deleted, and the reconstruction of the link in the all-optical network is realized through interaction with the controller.

[0029] In a possible implementation, the port connection relationship is used to indicate the number of connection ports and connection port identifiers between each switch group; the optical space switching device reconstructs and configures the optical path connection between the switch groups according to the link allocation information, and the optical path connection relationship between the optical space switching device and each switch group is formed, which may include: the optical space switching device determines the optical path port connected to each switch group among the optical path ports included in the optical space switching device according to the number of connection ports and connection port identifiers between each switch group; the optical space switching device configures the optical path connection between the switch groups based on the optical path port connected to each switch group, and the optical path connection relationship between the optical space switching device and each switch group is formed.

[0030] Based on this possible implementation, the optical space switching device can determine the optical path port based on the number of connection ports indicated by the controller and the connection port, and establish an optical path connection relationship between the optical space switching device and each switch group.

[0031] In a possible implementation, the method further includes: the optical space switching device responds to a link deletion request sent by the controller to instruct the optical space switching device to delete the port connection relationship indicated by the target link allocation information, and based on the target link allocation information carried in the link deletion request, deletes the port connection relationship indicated by the target link allocation information.

[0032] Based on this possible implementation, when a link is not needed, the port connection relationship corresponding to the link can be deleted based on a link deletion request, thereby releasing network resources, improving network resource utilization, and relieving storage pressure on the port connection relationship.

[0033] In a possible implementation, the optical space switching device deleting the port connection relationship indicated by the target link allocation information may include: disconnecting the optical path port connected to each target switch group, and deleting the port connection relationship indicated by the target link allocation information.

[0034] Based on this possible implementation, when a link is not needed, the optical path port connected to the switch group can be disconnected and the locally stored port connection relationship can be deleted, thereby deleting the port connection relationship indicated by the target link allocation information and simplifying the system design.

[0035] In a third aspect, a link reconstruction method is proposed, which is applied to a cluster including an all-optical network and a controller, wherein the all-optical network includes an optical space switching device and a switch group, and the method includes: the controller obtains resource information indicating each switch group executing a service in the all-optical network and the computing power parameters of each switch group, and determines link allocation information indicating the port connection relationship between the switch groups executing the service according to the computing power parameters of each switch group; the computing power parameters are used to indicate the computing power of the switch group when executing the service; the optical space switching device receives the link allocation information from the controller, and reconstructs and configures the optical path connection between the switch groups according to the link allocation information to form an optical path connection relationship between the optical space switching device and each switch group.

[0036] Based on the method described in the third aspect, when the controller receives the resource information of the service, it allocates the port connections between the switch groups according to the resource information of the service to realize the network topology reconstruction configuration. There is no need to monitor the distribution of network traffic, which can reduce the cost required for switches and controllers. At the same time, the links in the network are reconstructed and configured before the service is executed, avoiding the data packet loss caused by link switching during the task execution, and the waste of network resources caused by waiting for the data flow on the link to be drained.

[0037] In one possible implementation, the cluster also includes a scheduler, and the method also includes: the scheduler responds to a service application request, performs task scheduling based on the service resources carried in the service application request and currently idle computing nodes in the all-optical network, obtains service resource information, and sends the service resource information to the controller.

[0038] Based on this possible implementation method, the scheduler performs task scheduling based on the service resources carried in the service application request and the currently idle computing nodes in the all-optical network, obtains service resource information, and realizes load balancing of computing nodes in the all-optical network.

[0039] In a possible implementation, the method further includes: the controller sends a routing configuration instruction to each switch group according to the link allocation information; the switch group responds to the routing configuration instruction, performs routing configuration according to the link allocation information, and establishes a routing path between each switch group.

[0040] Based on this possible implementation, the creation of a routing path can be implemented based on the instruction interaction between the switch group and the controller, which simplifies the system design and improves the efficiency of creating the routing path.

[0041] In a possible implementation, the routing configuration instruction includes an instance creation instruction and a port binding instruction; the method further includes: the switch group responds to the instance creation instruction sent by the controller, creates a virtual routing instance corresponding to the service, and returns instance creation confirmation information to the controller; the switch group responds to the port binding instruction sent by the controller, determines the connection port to be bound, binds the connection port to the virtual routing instance, and establishes a routing path between each switch group.

[0042] Based on this possible implementation method, the routing configuration instructions can be divided into instance creation instructions and port binding instructions and indicated to the switch group, so that the switch group can execute the creation and binding of routing paths according to different instructions, simplifying the system design and improving the accuracy and efficiency of routing path creation and binding.

[0043] In a possible implementation, the virtual routing instance includes a shortest virtual routing instance and a non-shortest virtual routing instance; the connection port includes a first type of connection port connected to a computing node, and a second type of connection port connected to other switch groups; the first type of connection port is bound to the non-shortest virtual routing instance, and the second type of connection port is bound to the shortest virtual routing instance.

[0044] Based on this possible implementation method, a virtual routing instance corresponding to the service is created in the switch group, and routing isolation is implemented based on the virtual routing instance to avoid the phenomenon of link grabbing between services, which helps to alleviate network congestion. In addition, by creating a non-shortest virtual routing instance containing a non-shortest routing path and a shortest routing path, different virtual routing instances are bound according to the port type. During the data transmission process, when the shortest routing path is congested, data transmission is carried out through the non-shortest routing path to alleviate the load balancing problem.

[0045] In a possible implementation, after establishing the routing paths between the switch groups, the method further includes: the switch group receives data corresponding to the service, and forwards the data to other switch groups other than the switch group that execute the service in the all-optical network through the virtual routing instance corresponding to the service.

[0046] Based on this possible implementation, data can be sent based on the established virtual routing instance, which simplifies system design and improves the efficiency and reliability of data transmission.

[0047] In a possible implementation, the virtual routing instance includes a shortest virtual routing instance and a non-shortest virtual routing instance, the shortest virtual routing instance includes a first path, the non-shortest virtual routing instance includes a first path and a second path, the first path represents that the switch group is directly connected to the remaining switch groups that perform the service, and the second path represents that the switch group is communicated and connected with the remaining switch groups that perform the service through an intermediate switch group; the shortest virtual routing instance is used to forward data to the remaining switch groups that perform the service in the all-optical network through the first path; the non-shortest virtual routing instance is used to forward data to the remaining switch groups that perform the service in the all-optical network through the second path when the traffic characteristic of the first path is greater than a preset traffic characteristic threshold; when the traffic characteristic of the first path is less than or equal to the preset traffic characteristic threshold, the data is forwarded to the remaining switch groups that perform the service in the all-optical network through the first path; the traffic characteristic is used to indicate the amount of transmission between switch groups when performing the service.

[0048] Based on this possible implementation, non-shortest virtual routes are deployed in the reconstructed network to reduce the congestion of the shortest path and solve the load balancing problem. At the same time, non-shortest virtual routes are established based on paths with smaller traffic characteristics to further reduce the load balancing problem.

[0049] In the fourth aspect, a link reconstruction device is proposed. The link reconstruction device can be a controller or a chip or system on chip in the controller, and can also be a functional module in the controller for implementing the method in the first aspect or any possible implementation of the first aspect. The link reconstruction device can implement the function performed by the controller in the above-mentioned first aspect or the possible implementation of the first aspect, and the function can be implemented by hardware executing the corresponding software. The hardware or software includes one or more modules corresponding to the above-mentioned functions. For example, the link reconstruction device may include an acquisition module, a processing module, and a sending module. Among them, the acquisition module is used to obtain resource information indicating each switch group executing a service in the all-optical network and the computing power parameters of each switch group; the computing power parameters are used to indicate the computing power of the switch group when executing the service. The processing module is used to determine the link allocation information according to the computing power parameters of each switch group; the link allocation information is used to indicate the port connection relationship between the switch groups executing the service. The sending module is used to send the link allocation information to the optical space switching device in the all-optical network.

[0050] Specifically, the relevant processing actions of the acquisition module, the processing module and the sending module, and their intended effects can be referred to in the first aspect or any possible implementation method of the first aspect, and will not be described in detail.

[0051] In the fifth aspect, a link reconstruction device is proposed, which can be a controller or a chip or system on chip in the controller. The link reconstruction device can implement the function performed by the controller in the above-mentioned first aspect or the possible implementation of the first aspect, and the function can be implemented by hardware. In one possible implementation, the link reconstruction device includes a processor and a communication interface. Among them, the processor and the communication interface are used to support the link reconstruction device to execute the method in the first aspect or any possible implementation of the first aspect. In another possible implementation, the link reconstruction device may also include a memory, which is used to store computer-executable instructions and data necessary for the link reconstruction device. When the link reconstruction device is running, the processor executes the computer-executable instructions stored in the memory, so that the link reconstruction device executes the method described in the above-mentioned first aspect or any possible implementation of the first aspect.

[0052] In the sixth aspect, a link reconstruction device is proposed, which can be an optical space switching device or a chip or system on chip in the optical space switching device, and can also be a functional module in the optical space switching device for implementing the method in the second aspect or any possible implementation of the second aspect. The link reconstruction device can implement the function performed by the optical space switching device in the above-mentioned second aspect or the possible implementation of the second aspect, and the function can be implemented by hardware executing the corresponding software. The hardware or software includes one or more modules corresponding to the above-mentioned functions. For example, the link reconstruction device may include a receiving module and a link processing module.

[0053] A receiving module, configured to receive link allocation information from a controller for indicating a port connection relationship between switch groups that execute the service; the resource information is used to indicate each switch group that executes the service in the all-optical network and a computing power parameter of each switch group, and the link allocation information is determined based on the computing power parameter of each switch group;

[0054] The link processing module is used to reconfigure the optical path connection between the switch groups according to the link allocation information to form an optical path connection relationship between the optical space switching device and each of the switch groups.

[0055] Specifically, the relevant processing actions of the receiving module and the link processing module, and their intended effects can be referred to in the second aspect or any possible implementation of the second aspect, and will not be described in detail.

[0056] In the seventh aspect, a link reconstruction device is proposed, which can be an optical space switching device or a chip or system on chip in the optical space switching device. The link reconstruction device can implement the function performed by the optical space switching device in the above-mentioned second aspect or the possible implementation of the second aspect, and the function can be implemented by hardware. In one possible implementation, the link reconstruction device includes a processor and a communication interface. Among them, the processor and the communication interface are used to support the link reconstruction device to execute the method in the second aspect or any possible implementation of the second aspect. In another possible implementation, the link reconstruction device may also include a memory, which is used to store computer-executable instructions and data necessary for the link reconstruction device. When the link reconstruction device is running, the processor executes the computer-executable instructions stored in the memory, so that the link reconstruction device executes the method described in the above-mentioned second aspect or any possible implementation of the second aspect.

[0057] In the eighth aspect, a cluster is proposed, which can also be replaced by a computer cluster or a computer network or a computer system without limitation. The cluster includes an all-optical network, a controller and a scheduler. The all-optical network includes an optical space switching device and a switch group. The cluster is used to execute the link reconstruction method in the above third aspect or any possible implementation method of the third aspect.

[0058] Based on the above scheme, the link reconstruction method, device and cluster of the embodiments of the present application, when the controller receives the resource information of the service, allocates the port connection between the switch groups according to the resource information of the service, and realizes the network topology reconstruction configuration. There is no need to monitor the distribution of network traffic, which can reduce the cost required for switches and controllers. At the same time, the network topology reconstruction configuration is performed before the service is executed to avoid data packet loss caused by link switching during task execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 This is a schematic diagram of the topological structure of dragonfly;

[0060] Figure 2 This is a schematic diagram of the mapping relationship adjustment of the OXC device;

[0061] Figure 3 is a flowchart of a link reconstruction method;

[0062] Figure 4 This is a schematic diagram of an application scenario of the link reconstruction method provided in an embodiment of the present application in an AI cluster network;

[0063] Figure 5 is a schematic diagram of the structure of the AI ​​cluster network 423 provided in an embodiment of the present application;

[0064] Figure 6 is a schematic diagram of the structure of the AI ​​cluster network 423 after link reconstruction provided in an embodiment of the present application;

[0065] Figure 7 It is a flowchart of a link reconstruction method provided in an embodiment of the present application;

[0066] Figure 8 is a schematic diagram of link allocation between switch groups provided in an embodiment of the present application;

[0067] Fig. 9 It is a flowchart of a method for determining the number of link allocations provided in an embodiment of the present application;

[0068] Fig.10 This is a schematic diagram of link allocation between switch groups based on link allocation ratios provided in an embodiment of the present application;

[0069] Fig.11 It is a flowchart of a method for determining the number of link allocations using remaining ports provided in an embodiment of the present application;

[0070] Fig.12 The embodiment of the present application provides Fig.10 Schematic diagram of link allocation between switch groups;

[0071] Fig.13 It is a schematic diagram of allocating remaining ports based on traffic characteristics provided by an embodiment of the present application;

[0072] Fig.14 It is a schematic diagram of link isolation of different tasks provided in an embodiment of the present application;

[0073] Fig.15 This is a schematic diagram of routing isolation provided in an embodiment of the present application;

[0074] Fig.16 Schematic diagram of the shortest routing path and the non-shortest routing path provided in the embodiment of the present application;

[0075] Fig.17 is a schematic structural diagram of a link reconstruction device 17 provided in a controller according to an embodiment of the present application;

[0076] Fig.18 It is a structural schematic diagram of a link reconstruction device 18 provided in an optical space switching device provided in an embodiment of the present application;

[0077] Fig.19 It is a schematic diagram of the structure of a cluster to which the link reconstruction method is applied provided in an embodiment of the present application;

[0078] Fig. 20 is a structural schematic diagram of a controller provided in an embodiment of the present application;

[0079] Fig.21 This is another structural diagram of the controller provided in an embodiment of the present application. DETAILED DESCRIPTION

[0080] Before describing the specific implementation methods, the relevant technical terms involved in this application are described:

[0081] All-optical network refers to a direct optical fiber communication network that directly completes all network communication functions at the optical layer, that is, directly performs random storage, transmission and exchange processing of signals in the optical domain, and replaces the electrical nodes of the existing network with optical nodes in the network, and is based on optical fiber. Optionally, all-optical networks can include but are not limited to high performance computing (HPC) networks and artificial intelligence (AI) cluster networks.

[0082] A high performance computing (HPC) network consists of hundreds or thousands of servers connected by a network. Each server acts as a computing node. The computing nodes in the HPC network work in parallel with each other to increase processing speed, thereby achieving high performance computing.

[0083] An artificial intelligence (AI) cluster network may include multiple graphic processing units (GPUs) connected through a network. Each CPU may serve as a computing node. Distributed training of AI models may be achieved through multiple GPUs, thereby improving the training speed of AI models.

[0084] In HPC networks and AI cluster networks, commonly used topologies include fat-tree topology, dragonfly topology, and torus topology. Figure 1 As shown, Figure 1 This is a schematic diagram of the topology of dragonfly. In the topology of dragonfly, multiple adjacent switches are fully interconnected to form a group, which can be called a switch group. The groups are fully interconnected to form the entire network topology. The switch configuration in the HPC network and the AI ​​cluster network is the same. Each switch contains three types of ports, which are used to connect the ports of the computing nodes in the network, the ports connecting the switches in the group, and the ports connecting the switches outside the group. The switches in the group are connected using shorter local links, and the switches across the group that are farther away are connected using global links. Dragonfly uses a fully interconnected network topology both within and between groups, which effectively reduces the network diameter. The communication between any two switches only involves three hops at most.

[0085] For example, Dragonfly includes 7 switch groups: group 1 to group 7. The switches in group 1 are all interconnected. Switch A is connected to switch B in group 1. Switch C in group 7 is connected to switch D in group 7. All switch groups are fully interconnected. Switch B in group 1 is connected to switch C in group 7. When switch A in group 1 needs to communicate with switch D in group 7, Figure 1 In the direction of the arrow, switch A in group 1 can send the data to be transmitted to switch B in group 1. Switch B in group 1 sends the received data to be transmitted by switch A to switch C in group 7. Switch C in group 7 then forwards the data to be transmitted by switch A to switch D in group 7.

[0086] In order to achieve flexible adjustment of the topology of the HPC network and the AI ​​cluster network, optical space switching devices can be added to the HPC network and the AI ​​cluster network, and the mapping relationship between ports (or port mapping relationship) can be configured through the optical space switching device, and the connection relationship between switches can be adjusted based on the mapping relationship. Among them, the optical space switching device includes but is not limited to a micro-electro-mechanical-system-based optical switch (MEMS-based optical switch, MBOS) or an optical cross-connect (OXC) device.

[0087] For example, if the optical space switching device is an OXC device, the OXC device is connected to four switches (SW1-SW4), and the connection relationship between the four switches (SW1-SW2) can be adjusted through the OXC device. Figure 2 As shown in Figure (a), the following mapping relationship is configured in the OXC device: SW1_port 1: SW2_port 5, SW1_port 2: SW3_port 9, SW1_port 3: SW3_port 10, SW1_port 4: SW4_port 13, SW2_port 6: SW3_port 11, SW2_port 7: SW4_port 14, SW2_port 8: SW4_port 15, SW3_port 12: SW4_port 16. Based on this mapping relationship, it can be realized that: port 1 in switch SW1 is connected to switch SW 2, port 2 in switch SW1 is connected to port 9 of switch SW3, port 3 in switch SW1 is connected to port 10 of switch SW3, port 4 in switch SW1 is connected to port 13 of switch SW4, port 6 in switch SW2 is connected to port 11 of switch SW3, port 7 in switch SW2 is connected to port 14 of switch SW4, port 8 in switch SW2 is connected to port 15 of switch SW4, and port 12 in switch SW3 is connected to port 16 of switch SW4. Subsequently, the original mapping relationship can be adjusted as needed, so that the adjusted mapping relationship is: SW1_port1: SW2_port5, SW1_port2: SW3_port9, SW1_port3: SW3_port8, SW1_port4: SW4_port13, SW2_port6: SW3_port11, SW2_port7: SW4_port14, SW3_port10: SW4_port15, SW3_port12: SW4_port16. In this way, the OXC device can Figure 2In (a), the number of connection lines between switch SW1 and switch SW2 is reduced to 1, and the number of connection lines between switch SW1 and switch SW2 is increased to 2, forming Figure 2 The connection relationship after mapping is shown in Figure (b).

[0088] Optionally, the mapping relationship configured in the OXC device can be adaptively adjusted according to the traffic distribution of the HPC network and the AI ​​cluster network. Figure 3 As shown, Figure 3 It is a flowchart of a link reconstruction method. During the execution of a service, the traffic distribution of the network is monitored, the monitored traffic distribution is stored, the stored historical traffic distribution and the current network traffic distribution are periodically read, and the traffic characteristics are summarized. The topology matrix is ​​calculated based on the traffic characteristics, and the port mapping relationship of the OXC is formed based on the topology matrix, and the shortest route is configured. The link is reconstructed through the port mapping relationship of the OXC and the shortest route to achieve on-demand allocation of network bandwidth. It should be understood that the link reconstruction described in this application may refer to disconnecting the original link between ports and connecting them to other ports. The link reconstruction can also be replaced by link switching or other names without limitation.

[0089] However, the monitoring of the traffic distribution of the network (or traffic monitoring) is complex and costly. Different tasks have different iteration cycles, and the monitoring and topology switching time is difficult to determine. In addition, historical traffic information needs to be stored additionally. It is necessary to add additional storage devices to the controller and optical space switching device used to control link switching, or to select a controller and optical space switching device with a larger storage space, which will increase the hardware cost of the controller and optical space switching device, and the monitored traffic may not be accurate. In addition, the above-mentioned link reconstruction process occurs during the execution of the service. The link is unavailable during the link reconstruction period, which affects the data flow of the service being transmitted. For example, directly disconnecting a link will cause the transmission of the data flow that is about to use this link to fail, resulting in packet loss. The common practice is to drain the data flow on the link before disconnecting the link until the link is temporarily idle and then start the link reconstruction. This emptying method may require a long waiting time, which affects the transmission of the service data flow and wastes network resources.

[0090] Based on this, an embodiment of the present application provides a link reconstruction method, which may include: obtaining resource information of the business, allocating port connections between each switch group according to the resource information of the business, and realizing network topology reconstruction configuration. There is no need to monitor the distribution of network traffic, which can reduce the cost required for switches and controllers. At the same time, the reconstruction configuration of the links in the network is performed before the business is executed, avoiding data packet loss caused by link switching during task execution, and waiting for the link to be empty, causing waste of network resources.

[0091] The link reconstruction method provided in the embodiment of the present application will be described in detail below in combination with application scenarios.

[0092] The link reconstruction method provided in the embodiment of the present application can be applied to all-optical networks such as AI cluster networks and HPC networks. The following takes the AI ​​cluster network with Dragonfly as an example to introduce the link reconstruction method provided in the embodiment of the present application. It should be understood that the link reconstruction method in other all-optical networks (such as HPC networks) can be executed with reference to it and will not be described in detail.

[0093] like Figure 4 As shown, Figure 4 Schematic diagram of an application scenario of the link reconstruction method provided in an embodiment of the present application in an AI cluster network, wherein the application scenario shown includes a user end 410, a server end 420 deployed with an AI cluster network 423, a controller 422, and a scheduler 421. Data is transmitted between the server end and the user end through a communication network.

[0094] The user terminal 410 is used to initiate a task application, such as an AI training task application, to the server 420. In some implementations, the user terminal 410 may be a computer, a mobile phone, a tablet computer, or an application deployed on a terminal, such as a program editing application, a development application, etc.

[0095] The server 420 is used to receive the task application initiated by the user terminal 410, call the resources in the AI ​​cluster network 423 to execute the task initiated by the user terminal 410, and feed back the task result to the user terminal 410. For example, in the AI ​​training scenario, the server 420 can receive an AI training task application from the user terminal 410, execute the AI ​​training task based on the AI ​​training task application to obtain a trained AI model, and return the trained AI model to the user terminal 410.

[0096] In some embodiments, the server 420 includes an application layer and a processing layer, wherein an AI cluster network 423 is deployed in the processing layer for performing AI training tasks. A communication interface and a calling interface are provided in the application layer. The application layer can establish a communication connection with the user terminal 410 through the communication interface, and the communication interface can be used to receive an AI training task application initiated by the user terminal 410, and return the trained AI model to the user terminal 410. The calling interface is used to pass the AI ​​training task to the processing layer, and to obtain the trained AI model obtained by the processing layer based on the AI ​​training task training. The controller 422 and the scheduler 421 can be set at the application layer, or at the processing layer, or the scheduler 421 can be set at the application layer and the controller 422 can be set at the processing layer.

[0097] In a possible implementation, switches, computing nodes, and optical space switching devices (such as OXC devices) are deployed in the AI ​​cluster network 423. The computing nodes are connected to the switches, and at least one computing node can be connected to one switch. Each switch is connected to the optical space switching device, and the optical space switching device is configured with a mapping relationship between the ports of the switch, which can be used to establish a link between switches.

[0098] The switches may be connected in a fully interconnected manner, with each switch connected to the remaining switches in the AI ​​cluster network 423; the switches may also be connected in a tree structure, with each switch connected to the switch corresponding to the parent node of the switch; the switches may also be connected in a switch group manner, with the switches in the AI ​​cluster network 423 divided into multiple switch groups, each switch group connected to the remaining switch groups in the AI ​​cluster network 423, and each switch in the switch group connected to the remaining switches in the switch group. Each switch group includes at least one switch. It is understandable that in the AI ​​cluster network 423, switches with similar spatial positions may be fully interconnected to form a switch group. It is understandable that each switch group is connected to the remaining switch groups in the AI ​​cluster network 423. It can be said that a master switch is set in the switch group, and the master switch in each switch group is connected to the master switches in the remaining switch groups in the AI ​​cluster network 423, wherein the master switch may be specified when the AI ​​cluster network 423 is created.

[0099] Among them, the computing node can have the functions of computing and providing training services for the AI ​​training tasks applied for by the user terminal 410. The computing node can be replaced by a computing device. For example, the computing node can be a server, a high-performance computer, a cloud server, etc., or a processor, such as a central processing unit (CPU), a GPU, and a neural network processing unit (NPU), etc., or the computing node can be a functional module in the server that implements computing functions and provides training services.

[0100] Among them, the optical space switching device is used to reconstruct the links between the switches performing AI training tasks according to the link allocation information sent by the controller 422; the switches performing AI training tasks respond to the routing configuration instructions to establish routing paths between the switches performing AI training tasks.

[0101] For example, the AI ​​cluster network 423 includes 8 switches, each switch is connected to 2 servers, and each server has 8 GPUs deployed, that is, each switch is connected to 16 computing nodes. Figure 5 As shown, Figure 5 4 is a schematic diagram of the structure of the AI ​​cluster network 423 provided in the embodiment of the present application, where each switch is connected to the other 7 switches through an OXC device. It should be noted that: Figure 5 The AI ​​cluster network 423 structure shown is only an exemplary description. In actual applications, the AI ​​cluster network 423 may include Figure 5 A greater or lesser number of switches in the

[0102] In some embodiments, the scheduler 421 and the controller 422 may be integrated into one device, for example, the scheduler 421 and the controller 422 may be integrated into a computing node in the server 420. In other embodiments, the scheduler 421 and the controller 422 may be coupled and connected, for example, the scheduler 421 may be set in a computing node in the server 420, and the controller 422 may be set in the computing node where the scheduler 421 is located in the server 420, and coupled and connected with the scheduler 421. For example, the scheduler 421 and the controller 422 may also be deployed in two different computing nodes in the server 420, respectively, and coupled and connected with the scheduler 421.

[0103] Among them, the scheduler 421 can be connected to the controller 422, and the scheduler 421 can be used to schedule computing resources, storage resources, network resources and other resources in the AI ​​cluster network 423, such as scheduling computing nodes, switches, etc. in the AI ​​cluster network 423, and can also be used to send the scheduling results / scheduling information to the controller 422, so that the controller 422 can perform network control on the AI ​​cluster network 423 according to the scheduling results / scheduling information. For example, in the AI ​​training scenario, the scheduler 421 can perform task scheduling according to the number of computing nodes required for the AI ​​training task applied by the user end 410, determine the computing nodes, switches and the number of computing nodes assigned to each switch for executing the AI ​​training task, and send the service identifier of the AI ​​training task, the switch and the number of computing nodes assigned to each switch to the controller 422.

[0104] Among them, the controller 422 is connected to the scheduler 421, the switch, and the optical space switching device respectively. The controller 422 can be used to perform network control on the AI ​​cluster network 423 based on the scheduling result / scheduling information, such as controlling the link allocation, network topology, etc. in the AI ​​cluster network 423, and can also be used to send configuration information to the nodes in the AI ​​cluster network 423 (such as optical space switching devices and / or switches) to control the operation of the nodes in the AI ​​cluster network 423. For example, taking the AI ​​training scenario as an example, after receiving the switch sent by the scheduler 421 and the number of computing nodes allocated to each switch, the controller 422 performs link allocation based on the scheduling information sent by the scheduler 421, obtains the number of link allocations between the switches connecting the computing nodes that perform the AI ​​training task, and the connection ports between the switches connecting the computing nodes that perform the AI ​​training task, forms link allocation information, sends the link allocation information to the optical space switching device in the AI ​​cluster network 423, and sends routing configuration instructions to the switches connected to the computing nodes that perform the AI ​​training task, and establishes a routing path for the switches connecting the computing nodes that perform the AI ​​training task.

[0105] Still Figure 5 Taking the AI ​​cluster network 423 structure provided as an example, when the user end 410 applies to the server end 420 for an AI training task XX1, the number of computing nodes required is 32, and the computing nodes in the current AI cluster network 423 are all idle, the scheduler 421 determines through task scheduling that the four servers S0-S3 under the switches A1 and A2 are allocated to the AI ​​training task XX1 applied by the user end 410, each server uses 8 computing nodes, and the number of computing nodes allocated to the switches A1 and A2 is 16. The scheduler 421 allocates the switches A1, A2, The controller 422 determines that the number of links allocated between switches A1 and A2 is 14 based on the number of computing nodes allocated to switches A1 and A2, and sends the connection ports K1 to K14 between switches A1 and A2 to the optical space switching device. The optical space switching device directly connects the connection ports K1 to K14 of switch A1 and the connection ports K14 to K1 of A2, forming a Figure 6 The AI ​​cluster network 423 structure shown associates the connection ports K1 to K14 of switches A1 and A2 with the AI ​​training task XX1.

[0106] The following is based on Figure 4 The application scenario provided is used to describe in detail the link reconstruction method provided in the embodiment of the present application. The link reconstruction method can be performed by Figure 4The controller 422 in the embodiment may also be executed by a computing device with data processing capabilities or by other functional modules having the functions executed by the controller 422, without limitation. The following takes the link reconstruction method executed by the controller 422 as an example for description. Figure 7 As shown, Figure 7 is a flow chart of a link reconstruction method provided in an embodiment of the present application, wherein the link reconstruction method shown at least includes steps S110 to S140:

[0107] Step S110: The controller obtains resource information of the service.

[0108] In some implementations, the service refers to the service applied for by the user terminal 410, which includes but is not limited to AI training service, data computing service, operations planning service, image recognition service, image rendering service, etc.

[0109] The resource information can be used to indicate the switch groups that execute services in the all-optical network and the computing power parameters of each switch group. Figure 4 In the AI ​​cluster network 423 in the application scenario shown, the switch group includes at least one switch. It is understandable that in an all-optical network, switches with similar spatial positions can be fully interconnected to form a switch group. The computing power parameter of the switch group can be a computing performance parameter of a computing node connected to the switch group, and the computing performance parameter can calculate the amount of data processed by the node per unit time, such as 32 bits (bit), 256 bits. Alternatively, the computing power parameter of the switch group can be the number of computing nodes occupied by the switch group when executing the business.

[0110] In some embodiments, the resource information of the service may be a data list, for example, the resource information of the service may be a list generated by associating the identification information of each switch group executing the service in the all-optical network with the computing power parameters of each switch group. In some other embodiments, the resource information of the service may be a string of data consisting of multiple fields, wherein each field records the identification information of the switch group and the computing power parameters of the switch group. Among them, the identification information of each switch group executing the service may be the serial number, name, device serial number of the switch group, or the network address of the computing node connected to the switch group. For example, taking the computing power parameter as the number of computing nodes occupied by the switch group when executing the service as an example, the resource information of the service ZY_XXX1:EX_S1:8;EX_S2:16;EX_S3:8, representing the switch groups that execute the service identified as XXX1 are S1, S2 and S3, and the number of computing nodes occupied by the switch groups S1, S2 and S3 when executing the service identified as XXX1 are 8, 16 and 8 respectively.

[0111] In some implementations, the resource information of the service may be obtained by a scheduler performing task scheduling and sent by the scheduler to the controller, or may be sent by other devices to the controller.

[0112] In a possible implementation, the scheduler, in response to the service application trigger request, performs job scheduling and resource allocation, forms service resource information, and sends it to the controller.

[0113] The service application trigger request is used to request the all-optical network to execute the service. The service application request carries the service resources of the service, wherein the service resources are used to represent the computing resources required to execute the service.

[0114] For example, the scheduler can perform task scheduling based on the service resources carried in the service application request and the currently idle computing nodes in the all-optical network according to the existing task scheduling method, obtain the computing nodes executing the service, the number of computing nodes, and the switch groups connected to the computing nodes, obtain the switch groups executing the service and the number of computing nodes occupied by the switch groups when executing the service according to the number of computing nodes and the switch groups connected to the computing nodes, use the switch groups executing the service and the number of computing nodes occupied by the switch groups when executing the service as the resource information of the service, and send a link application request to the controller based on the resource information of the service and the service identifier of the service. The controller responds to the link application request sent by the scheduler, obtains the resource information of the service, and executes steps S110 to S140 to achieve link reconstruction.

[0115] The link application request carries service resource information, which is used to instruct the controller to perform link allocation based on the service resource information carried in the link application request.

[0116] In a further implementation, the service application trigger request also includes a topology type. The scheduler sends a link application request to the controller based on the resource information of the service, the service identifier of the service, and the topology type carried in the service application trigger request. The controller responds to the link application request, obtains the topology type, the resource information of the service, and the service identifier of the service carried in the link application request, executes steps S120 to S130 to determine the link allocation information between the switch groups executing the service according to the resource information of the service and the topology type, associates the link allocation information with the service identifier of the service, sends the link allocation information to the optical space switching device, and establishes an optical path connection relationship matching the topology type between the switch groups executing the service. Among them, the topology type includes but is not limited to a mesh topology, a ring topology, a cross topology, a torus topology, etc.

[0117] Step S120, determining link allocation information according to the computing power parameters of each switch group.

[0118] The link allocation information is used to indicate the port connection relationship between the switch groups that execute the service. The port connection relationship can be used to characterize the connection ports between the switch groups that execute the service and the connection relationship between the connection ports. The connection port can refer to a port that supports establishing a connection or link with other ports, and the link established corresponding to the connection port supports the data flow of the transmission service.

[0119] In some embodiments, the link allocation information includes a port connection relationship, and the link allocation information can directly indicate the port connection relationship between the switch groups that perform the service. Optionally, the link allocation information includes the connection ports between the switch groups that perform the service and the connection relationship between the connection ports in a list form. For example, as shown in Table 1 below, the link allocation information is in a tabular form. As shown in Table 1, port 1 of switch group SW1 is connected to port 1 of switch group SW2, and port 2 of switch group SW1 is connected to port 2 of switch group SW2.

[0120] Table 1 Link allocation information

[0121] SW2_Port1 SW2_Port2 SW_PORT1 connect —— SW_PORT2 —— connect

[0122] It should be noted that the link allocation information shown in Table 1 is an exemplary description and does not constitute a limitation on the link allocation information provided in the embodiments of the present application. In actual applications, when the list includes the connection ports between the switch groups that execute the business and the connection relationship between the connection ports, the connection relationship between the connection ports can be represented in the form of a connection identifier. For example, when the connection identifier of the connection port is 1, it indicates that the two connection ports are connected. When the connection identifier of the connection port is 0, it indicates that the two connection ports are not connected.

[0123] In some other embodiments, the link allocation information may indirectly indicate the port connection relationship. For example, the link allocation information may include information for determining the port connection relationship. Further, the controller may determine the connection ports between the switch groups executing the service and the connection relationship between the connection ports based on at least the information included in the link allocation information to obtain the port connection relationship.

[0124] Optionally, the link allocation information includes the link allocation quantity of the switch group executing the service. The controller can obtain the connection ports between the switch groups executing the service and the connection relationship between the connection ports according to the link allocation quantity of the switch group executing the service and the currently idle connection ports of the switch group.

[0125] For example, the controller determines the number of target ports that need to be allocated to the switch group based on the number of links allocated to the switch group that executes the service, allocates connection ports of the target number from the currently idle connection ports of the switch group based on the number of target ports that need to be allocated to the switch group, and allocates connection ports of the switch group that connect to other switch groups among all switch groups that execute the service except the switch group based on the number of links established between the switch group and other switch groups among all switch groups that execute the service except the switch group, and obtains the connection ports between the switch groups that execute the service and the connection relationship between the connection ports based on the connection ports between each switch group and each other switch group.

[0126] Among them, when allocating connection ports of the switch group connected to other switch groups except the switch group among all switch groups executing the service according to the number of links established between the switch group and other switch groups except the switch group among all switch groups executing the service, the allocation can be performed in the order of the number of links. For example, when the number of links between the switch group 1 and the switch group 2 is 4, the number of links between the switch group 1 and the switch group 3 is 6, and the currently idle connection ports of the switch group 1 are K1 to K12, the connection ports between the switch group 1 and the switch group 3 are allocated first in the order from large to small in the number of links, and K1 to K6 are used as the connection ports between the switch group 1 and the switch group 3. Then, the connection ports between the switch group 1 and the switch group 2 are allocated, and K7 to K10 are used as the connection ports between the switch group 1 and the switch group 2.

[0127] Among them, when allocating connection ports of the switch group connected to other switch groups except the switch group among all switch groups executing the service according to the number of links established between the switch group and other switch groups except the switch group among all switch groups executing the service, the allocation can also be performed according to the order of the identification information of the switch groups. For example, when the number of links between switch group 1 and switch group 2 is 4, the number of links between switch group 1 and switch group 3 is 6, and the currently idle connection ports of switch group 1 are K1 to K12, the connection ports between switch group 1 and switch group 2 are first allocated according to the order of the identification information from low to high, and K1 to K3 are used as the connection ports between switch group 1 and switch group 2, and then the connection ports between switch group 1 and switch group 2 are allocated, and K4 to K10 are used as the connection ports between switch group 1 and switch group 3.

[0128] The number of links allocated to the switch group executing the service may refer to the total number of links established between the switch group and other switch groups except the switch group in all switch groups executing the service. The idle connection port may refer to a port currently unallocated in the switch group.

[0129] For example, Figure 5 Taking the AI ​​cluster network 423 shown as an example, when the switch groups executing the service are switch group A1, switch group A2 and switch group A3, and the link allocation numbers of switch group A1, switch group A2 and switch group A3 are 8, 8 and 8 respectively, the sum of the link numbers between switch group A1 and switch group A2 and the link numbers between switch group A1 and switch group A3 is 8, the sum of the link numbers between switch group A2 and switch group A1 and the link numbers between switch group A2 and switch group A3 is 8, and the sum of the link numbers between switch group A3 and switch group A1 and the link numbers between switch group A3 and switch group A2 is 8. Figure 8 As shown, when the connection ports of switch group A1, switch group A2 and switch group A3 are all idle, switch group A1 allocates four connection ports K1~K4 to connect with connection ports K11~K14 of switch group A2, and allocates another four connection ports K5~K8 of switch group A1 to connect with connection ports K1~K4 of switch group A3, and switch group A2 allocates four connection interfaces K5~K8 to connect with connection interfaces K11~K14 of switch group A3.

[0130] In some implementations, the computing power parameters of each switch group may be input into a preset allocation model for link allocation to obtain link allocation information. The preset allocation model may be a mathematical model, such as a linear function or a nonlinear function. Alternatively, the preset allocation model may also be a model based on machine learning or a neural network, without limitation.

[0131] In some other implementations, link allocation may be performed according to the computing power parameter ratio of each switch group to obtain link allocation information.

[0132] In which, when the computing power parameter is a computing performance parameter of a computing node connected to the switch group, the computing power parameter ratio can be replaced by a computing performance parameter ratio.

[0133] For example, taking the switch group executing the service as switch A1, switch A2, and switch A3 as an example, when the computing power parameter of switch A1 is 64 bits, the computing power parameter of switch A2 is 128 bits, and the computing power parameter of switch A3 is 64 bits, link allocation is started with switch A1. Since the computing power parameter ratio among switch A1, switch A2, and switch A3 is 1:2:1, the number of links allocated between switch A1 and switch A2 and the number of links allocated between switch A1 and switch A3 satisfy the ratio relationship of 2:1. When the number of idle connection ports of switch A2 is currently 8, the number of links allocated between switch A1 and switch A2 is 4, and the number of links allocated between switch A1 and switch A3 is 2.

[0134] Among them, when the computing power parameter is the number of computing nodes occupied by the switch group when executing the business, the computing power parameter ratio can be replaced by the ratio of the number of computing nodes. When determining the number of links to be allocated between the switch groups, the number of links to be allocated between the switch groups can be allocated according to the ratio of the number of computing nodes occupied by each switch group when executing the business.

[0135] Specifically, the implementation process of allocating the number of links between the switch groups based on the proportion of the number of computing nodes occupied by each switch group when executing services can refer to the following Fig. 9 The corresponding embodiments.

[0136] Step S130: The controller sends link allocation information to an optical space switching device in the all-optical network, and the optical space switching device receives the link allocation information from the controller.

[0137] Step S140: the optical space switching device reconfigures the optical path connections between the switch groups according to the link allocation information, so as to form an optical path connection relationship between the optical space switching device and each switch group.

[0138] In some embodiments, the optical space switching device includes multiple ports and an interface for receiving control instructions, each port is connected to a switch group in the all-optical network, and the optical space switching device adjusts which ports are interconnected or disconnects which interconnected ports according to the received instructions, thereby achieving adjustment of the optical path connection relationship between the optical space switching device and each switch group.

[0139] Among them, one or more optical space switching devices may be deployed in the all-optical network.

[0140] In a possible implementation manner of the above embodiment, when a plurality of optical space switching devices are deployed in the all-optical network, each of the optical space switching devices are interconnected.

[0141] In another possible implementation manner of the above embodiment, when an optical space switching device is deployed in the all-optical network, the uplink ports of the switch groups in the all-optical network are connected to the downlink ports of the optical space switching device.

[0142] In another possible implementation of the above embodiment, when two or more optical space switching devices are deployed in the all-optical network, the switch groups in the all-optical network can be divided into different combinations, each combination is assigned an optical space switching device, and the uplink port of the switch group in each combination is connected to the downlink port of the optical space switching device in the combination. The number of combinations is equal to the number of optical space switching devices deployed in the all-optical network.

[0143] In some implementations, the optical space switching device is connected to a controller, receives link configuration information sent by the controller, and reconfigures the optical path connection between the switch groups according to the port connection relationship executed according to the link configuration information, thereby forming an optical path connection relationship between the optical space switching device and each switch group.

[0144] Specifically, the optical space switching device determines the connection ports between the switch groups that execute the services and the connection relationships between the connection ports according to the port connection relationships, determines the optical path ports connected to each switch group according to the connection ports between the switch groups that execute the services and the connection relationships between the connection ports, and determines the connection relationships between the optical path ports, and forms an optical path connection relationship between the optical space switching device and each switch group based on the optical path ports connected to each switch group according to the connection relationships between the optical path ports.

[0145] Among them, the connection relationship between the optical path ports is used to indicate which optical path ports connected to each switch group are interconnected. For example, taking the switch groups that execute the business as switch groups A1 and A2, when the port connection relationship indicates that the connection port KA0 of switch group A1 is connected to the connection port KB4 of switch group A2, the optical space switching device will connect the optical path port connected to the connection port KA0 of switch group A1 and the optical path port connected to the connection port KB4 of switch group A2.

[0146] Further optionally, after link reconstruction is implemented, the computing nodes executing the service process the service according to the node mapping relationship, and perform data transmission through the connection ports of each switch group executing the service.

[0147] Among them, the node mapping relationship is used to characterize the processing logic relationship between the computing nodes that execute the business. The controller establishes a node mapping relationship (or mapping relationship) based on the computing nodes connected to each switch group that executes the business in the resource information of the business. In this way, when there are multiple computing nodes processing the business, the processing logic relationship between the computing nodes that execute the business can be represented by the node mapping relationship to achieve distributed processing of the business and improve business processing efficiency.

[0148] The link reconstruction method provided in the embodiment of the present application is that when the controller receives the resource information of the service, it allocates the port connection between the switch groups according to the resource information of the service to realize the network topology reconstruction configuration. There is no need to monitor the distribution of network traffic, which can reduce the cost required for switches and controllers. At the same time, the reconstruction configuration of the links in the network is performed before the service is executed, avoiding the data packet loss caused by link switching during the task execution process, and the network resource waste caused by waiting for the data flow on the link to be drained.

[0149] In some implementations, the port connection relationship includes the number of connection ports and connection port identifiers between each switch group, and the number of connection ports and connection port identifiers between each switch group are used to represent the connection ports between the switch groups that execute the services and the connection relationship between the connection ports. For example, the port connection relationship: EX_S1:{EX_S2:4(K0~K3); EX_S3:4(K4~K7)}; EX_S2:{EX_S1:4(K0~K3); EX_S3:2(K0~K1)}; EX_S3{EX_S1:4(K0~K3); EX_S2:2(K4~K5)}, which indicates that the connection ports K0~K3 of the switch group EX-S1 are connected to the connection ports K0~K3 of the switch group EX_S2, the connection ports K4~K7 of the switch group EX_S2 are connected to the connection ports K0~K3 of the switch group EX_S3, and the connection ports K0~K1 of the switch group EX_S2 are connected to the connection ports K4~K5 of the switch group EX_S3.

[0150] The number of connection ports between the switch groups is the same as the number of link allocations between the switch groups.

[0151] In a possible implementation, the link allocation quantity between the switch groups executing the business is allocated by proportion of the computing power parameters between the switch groups, and a port connection relationship is established according to the link allocation quantity between the switch groups and the port information of each switch group to form link allocation information.

[0152] The port information includes currently idle connection ports of each switch group and connection port identifiers of currently idle connection ports.

[0153] In the above implementation, for any two switch groups, the number of connection ports between the two switch groups and the connection port identifiers of the connection ports between the two switch groups are determined according to the number of link allocations between the two switch groups, the currently idle connection ports of each switch group, and the connection port identifiers of the currently idle connection ports, thereby forming a port connection relationship between the two switch groups.

[0154] For example, taking the link allocation number between switch group 1 and switch group 2 as 4, when connection ports KA0~KA7 and KB0~KB7 in switch group 1 are currently idle, connection ports KA0~KA7 in switch group 2 have been allocated, and KB0~KB7 are currently idle, when the number of connection ports between switch group 1 and switch group 2 is 4, connection port KA0 in switch group 1 is connected to connection port KB0 in switch group 2, connection port KA1 in switch group 1 is connected to connection port KB1 in switch group 2, connection port KA2 in switch group 1 is connected to connection port KB2 in switch group 2, and connection port KA2 in switch group 1 is connected to connection port KB2 in switch group 2.

[0155] In the above implementation, the controller can obtain the maximum link allocation number of each switch group that executes the service, and allocate the link number according to the maximum link allocation number of each switch group that executes the service and the computing power parameter ratio between the switch groups that execute the service, so as to obtain the link allocation number between the switch groups.

[0156] Among them, the maximum link allocation number of the switch group represents the maximum value of the number of links that the switch group can connect to other switch groups. The maximum link allocation number of the switch group can be determined by the computing power parameter of the switch group. For example, in the case where the computing power parameter is the computing performance parameter of the computing node connected to the switch group, the mapping relationship between the preset computing performance parameter and the number of links can be queried according to the computing performance parameter to obtain the maximum link allocation number of the switch group. In the case where the computing power parameter is the number of computing nodes occupied by the switch group when executing the business, the number of computing nodes occupied by the switch group when executing the business can be determined as the maximum link allocation number of the switch group.

[0157] It is understandable that for each switch group, the sum of the link allocation numbers between the switch group and the other switch groups executing services is less than or equal to the maximum link allocation number of the switch group. In this way, the rationality of link allocation between each switch group can be ensured.

[0158] Among them, when the computing power parameter ratio is the computing performance parameter ratio, for each switch group executing the business, the computing performance parameter ratio between the remaining switch groups is multiplied by the maximum link allocation number of the switch group, and the product is rounded down to obtain the link allocation number between the switch group and the remaining switch groups.

[0159] For example, when the computing performance parameter ratio among switches A1, A2, and A3 is 1:2:1, when the maximum link allocation number of switch group A1 is 8, the maximum link allocation number of switch group A1 is set according to the computing performance parameter ratio of 2:1 between switches A2 and A3, and 4 links are allocated to switch group A2, and 2 links are allocated to switch group A3, that is, the link allocation number between switches A1 and A2 is 4, and the link allocation number between switches A1 and A3 is 2.

[0160] Among them, when the computing power ratio parameter is the ratio of the number of computing nodes, for each switch group executing the business, the ratio of the number of computing nodes between the remaining switch groups is multiplied by the maximum link allocation number of the switch group, and the product is rounded down to obtain the link allocation number between the switch group and the remaining switch groups.

[0161] When the computing power ratio parameter is the ratio of the number of computing nodes, in order to achieve a balance between the number of link allocations between the switch groups, when allocating the number of links according to the maximum number of link allocations of each switch group executing the service and the computing power parameter ratio between the switch groups executing the service, the link allocation can be started from the switch group occupying the smallest number of computing nodes, such as Fig. 9 As shown, Fig. 9 1 is a flow chart of a method for determining the number of link allocations provided in an embodiment of the present application. The method for determining the number of link allocations shown includes steps S121A to S123A:

[0162] Step S121A, selecting the first switch group from the switch groups executing the service in descending order according to the number of computing nodes occupied by each switch group when executing the service.

[0163] It can be understood that the first switch group selected for the first time is the switch group with the lowest number of computing nodes occupied when executing services, and the first switch group selected for the second time is the switch group with the second lowest number of computing nodes occupied when executing services.

[0164] Among them, when the number of computing nodes occupied by each switch group when executing the service is different, the first switch group is selected from the switch groups executing the service in order from low to high according to the number of computing nodes occupied by each switch group when executing the service.

[0165] Among them, when the switch groups that execute the task have the same number of computing nodes, and for the same order, there are switch groups that occupy the same number of computing nodes when executing the business, any one of the switch groups that occupy the same number of computing nodes when executing the business can be used as the first switch group currently selected, and the remaining switch groups in the switch groups that occupy the same number of computing nodes when executing the business can be used as the first switch group selected next time.

[0166] Exemplarily, taking switch groups A1, A2, A3, and A4 as examples, when the number of computing nodes occupied by switch groups A1, A2, A3, and A4 when executing services respectively is 16, 8, 8, and 4 respectively, switch group A4 is selected as the first switch group for the first time in order from low to high according to the number of computing nodes occupied by each switch group when executing services, and the number of link allocations between switch group A4 and switch groups A1, A2, and A3 are determined respectively according to steps S122A to S123A; since the number of computing nodes occupied by switch groups A2 and A3 when executing services is the same, when the first switch group is selected for the second time, any one of switch groups A2 and A3 can be used as the first switch group, for example, switch group A2 is used as the first switch group selected for the second time, and switch group A3 is used as the first switch group selected for the third time.

[0167] Step S122A, determining the link allocation ratio between the switch groups according to the ratio between the number of computing nodes occupied by the first switch group when executing services and the number of computing nodes occupied by the remaining switch groups when executing services.

[0168] The link allocation ratio is used to indicate the ratio of the number of computing nodes between the switch group and the remaining switch groups except the switch group in the switch group executing the service.

[0169] In a possible implementation of step S122A, the ratio between the number of computing nodes occupied by the first switch group when executing services and the number of computing nodes occupied by the remaining switch groups when executing services can be determined as the link allocation ratio between the first switch group and the remaining switch groups, and the link allocation ratio between each first switch group and the remaining switch groups is calculated in turn to obtain the link allocation ratio between the switch groups.

[0170] Step S123A, obtaining the link allocation quantity between each switch group according to the link allocation ratio between each switch group and the maximum link allocation quantity of each switch group.

[0171] The maximum number of links allocated to a switch group is equal to the number of computing nodes occupied by the switch group when executing a service. For example, when the number of computing nodes occupied by the switch group when executing a service is 16, the maximum number of links allocated to the switch group is 16, that is, the switch group can allocate up to 16 links to the remaining switch groups executing services.

[0172] In some implementations of step S123A, for each switch group that executes the service, the maximum link allocation number of the switch group is allocated to the remaining switch groups that execute the service according to the product of the link allocation ratio between the switch group and the remaining switch groups that execute the service and the maximum link allocation number of the switch group, so as to obtain the link allocation number between the switch group and the remaining switch groups that execute the service.

[0173] In a possible implementation of the above embodiment, in order to ensure that the sum of the link allocation quantities between the switch group and the remaining switch groups that perform services does not exceed the maximum link allocation quantity of the switch group, when performing link allocation based on the link allocation ratio between the switch group and the remaining switch groups that perform services, the product of the link allocation ratio between the switch group and the remaining switch groups that perform services and the maximum link allocation quantity of the switch group is rounded down to an integer, which is used as the link allocation quantity between the switch group and the remaining switch groups that perform services.

[0174] In another possible implementation of the above embodiment, the second switch group can be selected from the remaining switch groups executing the service in sequence, and the maximum link allocation quantity of the switch group is allocated to each second switch group according to the link allocation ratio between the switch group and each second switch group, so as to obtain the link allocation quantity between the switch group and each second switch group. The sum of the link allocation quantities between the switch group and each second switch group is less than or equal to the maximum link allocation quantity of the switch group.

[0175] Furthermore, the second switch group may be selected in order of the number of computing nodes occupied by the remaining switch groups executing the service from low to high. In an optional implementation, the second switch group may be randomly selected from the remaining switch groups executing the service.

[0176] For example, the switch groups executing the service are group A1, group A2 and group A3. The number of computing nodes occupied by group A1, group A2 and group A3 when executing the service are 16, 8 and 8 respectively. The link allocation starts from group A2. Group A2 is used as the first switch group selected for the first time. The maximum number of links allocated to group A2 is 8, and the link allocation ratio between group A1 and group A3 is 2:1. Therefore, according to the link allocation ratio of 2:1, the 8 links of group A2 are allocated to group A1 and group A2 respectively. Fig.10 As shown in Figure (a), there are 4 links between Group A1 and Group A2, and 2 links between Group A2 and Group A3. Then, Group A3 is selected as the first switch group for the second time. The maximum number of links allocated to Group A3 is 8, and the link allocation ratio between Group A1 and Group A2 is 2:1. Since the links between Group A3 and Group A2 have been allocated, the remaining 4 links of Group A3 are connected to Group A1 according to the link allocation ratio between Group A1 and Group A2, as shown in Fig.10 As shown in Figure (b), the number of links allocated between group A1 and group A2 is 4, the number of links allocated between group A1 and group A3 is 4, and the number of links allocated between group A2 and group A3 is 2.

[0177] based on Fig. 9 In the provided embodiment, the controller allocates the number of link allocations between switch groups according to the distribution ratio of the number of computing nodes occupied by the switches during business operations. There is no need to monitor the traffic in the network, thus avoiding the cost of traffic monitoring. At the same time, according to the distribution ratio of the number of computing nodes occupied by different switch groups during business operations, it is possible to improve the utilization rate of network links while achieving balanced link allocation.

[0178] Considering link allocation, since link allocation is performed according to the maximum link allocation number of the switch group and the link allocation ratio between the switch groups, the sum of the link allocation numbers between the switch group and other switch groups is less than the maximum link allocation number of the switch group. This will result in unused connection ports (remaining connection ports) in the connection ports allocated to the services by the switch group. Based on this, in some embodiments, in order to maximize the network link utilization, after link allocation is performed according to the maximum link allocation number of the switch group and the link allocation ratio between the switch groups, when there are switch groups with remaining connection ports, the remaining connection ports are utilized by increasing the links between the switch groups with the remaining connection ports to further improve the network link utilization.

[0179] Specifically, Fig.11 As shown, Fig.11 1 is a flow chart of a method for determining the number of link allocations using the remaining ports provided in an embodiment of the present application. The method for determining the number of link allocations using the remaining ports shown includes steps S121B to S125B:

[0180] Step S121B, selecting the first switch group from the switch groups executing the service in order from low to high according to the number of computing nodes occupied by each switch group when executing the service.

[0181] In some implementations, the first switch group may be selected from the switch groups executing the service with reference to the above step S121A, which will not be described in detail in this embodiment of the present application.

[0182] Step S122B, determining the link allocation ratio between the switch groups according to the ratio between the number of computing nodes occupied by the first switch group when executing services and the number of computing nodes occupied by the remaining switch groups when executing services.

[0183] In some implementations, the link allocation ratio between the switch groups may be determined with reference to the above step S122A, which is not described in detail in this embodiment of the present application.

[0184] Step S123B, obtaining the initial link allocation quantity between each switch group according to the link allocation ratio between each switch group and the maximum link allocation quantity of each switch group.

[0185] In some implementations, the initial link allocation quantity between each switch group may be obtained by referring to the above step S123A, which is not described in detail in this embodiment of the present application.

[0186] Step S124B, if the sum of the initial link allocation quantities between the switch group and each second switch group is equal to the maximum link allocation quantity of the switch group, then the initial link allocation quantity between the switch group and each second switch group is determined as the link allocation quantity between the switch group and each second switch group.

[0187] In a possible implementation of step S124B, if the sum of the initial link allocation quantities between the switch group and each second switch group is equal to the maximum link allocation quantity of the switch group, indicating that there are no remaining connection ports for use of the connection ports allocated to the services by the switch group, then the initial link allocation quantities between the switch group and each second switch group are not adjusted.

[0188] Exemplarily, taking the switch groups executing the business as group A1 and group A2 as an example, if the number of computing nodes occupied by group A1 and group A2 when executing the business are 8 and 8 respectively, then the maximum number of link allocations between group A1 and group A2 is 8. According to the above steps 121B to 123B, the initial number of link allocations between group A1 and group A2 is 8. Since the initial number of link allocations between group A1 and group A2 is equal to the maximum number of link allocations between group A1 and group A2, the initial number of link allocations between group A1 and group A2 is determined as the number of link allocations between group A1 and group A2.

[0189] Step S125B, after step S123B, if the sum of the initial link allocation quantities between the switch group and each second switch group is less than the maximum link allocation quantity of the switch group, the initial link allocation quantity between the switch group and the third switch group is adjusted to obtain the link allocation quantity between the switch group and each second switch group.

[0190] The third switch group is included in the second switch group whose link allocation number is less than the maximum link allocation number.

[0191] In a possible implementation of step S125B, when the sum of the initial link allocation quantities between the switch group and each second switch group is less than the maximum link allocation quantity of the switch group, the switch group with the least number of computing nodes occupied when executing the task can be selected from the second switch groups whose link allocation quantity is less than the maximum link allocation quantity as the third switch group, and the initial link allocation quantity between the switch group and the third switch group is increased according to the maximum link allocation quantity of the third switch group and the maximum link allocation quantity of the switch group, and the link allocation quantity between the switch group and each second switch group is obtained based on the adjusted initial link allocation quantity between the switch group and the third switch group and the initial link allocation quantity between the switch group and the second switch group other than the third switch group.

[0192] For example, the switch groups executing the service are group A1, group A2 and group A3. The number of computing nodes occupied by group A1, group A2 and group A3 when executing the service are 16, 8 and 8 respectively. According to the above steps 121B to 123B, the links between group A1, group A2 and group A3 are allocated. Fig.12 In Figure (a), the initial number of links allocated between group A1 and group A2 is 4, the initial number of links allocated between group A1 and group A3 is 4, and the initial number of links allocated between group A2 and group A3 is 2; since among the switch groups connected to group A2, group A3 occupies the least number of computing nodes when executing services, and both groups A2 and A3 have two remaining connection ports unused, two links are added between group A2 and group A3, as shown in Figure 1. Fig.12 As shown in Figure (b), the number of links allocated between group A1 and group A2 is 4, the number of links allocated between group A1 and group A3 is 4, and the number of links allocated between group A2 and group A3 is 4.

[0193] In another possible implementation of step S125B, when there are two or more second switch groups in each second switch group whose link allocation number is less than the maximum link allocation number, the third switch group may be selected from the second switch groups in which the link allocation number is less than the maximum link allocation number according to the traffic characteristics between the switch group and each second switch group. The traffic characteristics are used to indicate the amount of data transmission between the switch groups when executing the service.

[0194] For example, according to the traffic characteristics between the switch group and each second switch group, the switch group with the largest traffic characteristics in the second switch group whose link allocation number is less than the maximum link allocation number can be determined as the third switch group.

[0195] For example, according to the traffic characteristics between the switch group and each second switch group, the switch group with the smallest traffic characteristics in the second switch group whose link allocation number is less than the maximum link allocation number can be determined as the third switch group.

[0196] For example, the switch groups executing the service are group A1, group A2 and group A3. The number of computing nodes occupied by group A1, group A2 and group A3 when executing the service are 16, 8 and 8 respectively. According to the above steps 121B to 123B, the links between group A1, group A2 and group A3 are allocated. Fig.13 In the figure (a), the initial link allocation number between group A1 and group A2 is 4, the initial link allocation number between group A1 and group A3 is 4, and the initial link allocation number between group A2 and group A3 is 2. When the traffic characteristics between group A1 and group A2 are 32, and the traffic characteristics between group A2 and group A3 are 64, if the switch group with the smallest traffic characteristics in the second switch group whose link allocation number is less than the maximum link allocation number is determined as the third switch group, then group A1 is used as the third switch group. Since the remaining two connection ports in group A2 and the remaining two connection ports in group A3 are not used, two links are added between group A2 and group A1, and two links are added between group A1 and group A3 according to the ratio of the number of computing nodes occupied by group A2 and group A3 when executing services, as shown in FIG. Fig.13 As shown in Figure (b), the number of links allocated between group A1 and group A2 is 6, the number of links allocated between group A1 and group A3 is 6, and the number of links allocated between group A2 and group A3 is 2. If the switch group with the largest traffic characteristics in the second switch group whose link allocation number is less than the maximum link allocation number is determined as the third switch group, then group A3 is used as the third switch group. Since there are two unused connection ports in both group A2 and group A3, two links are added between group A2 and group A3, as shown in Figure 1. Fig.13 As shown in Figure (c), the number of links allocated between group A1 and group A2 is 4, the number of links allocated between group A1 and group A3 is 4, and the number of links allocated between group A2 and group A3 is 4.

[0197] based on Fig.11 In the provided embodiment, when the controller allocates links according to the distribution ratio of the number of computing nodes occupied by the switch groups when executing services, when there are switch groups with remaining connection ports, the remaining connection ports are utilized by increasing the links between the switch groups with remaining connection ports to further improve the network link utilization rate; and according to the traffic characteristics between the switch groups, the switch groups that utilize the remaining connection ports are selected to further achieve traffic balance between the switch groups.

[0198] Considering that in the prior art, when link reconstruction is performed, switching links during task operation will seriously affect the data stream being transmitted, and directly disconnecting a link will cause the flow transmission that is about to use this link to fail, resulting in packet loss, or during the task execution, there are situations where different tasks occupy the same link, affecting the execution efficiency of the business. Based on this, in order to solve the problem of data packet loss and link preemption caused by link reconstruction, in some implementations, before the business is executed, the computing nodes, switch groups, and the number of computing nodes occupied by each switch group when executing the business are allocated to execute the business, and based on the switch groups that execute the business and the number of computing nodes occupied by each switch group when executing the business, link allocation is performed to obtain the number of link allocations between the switch groups that execute the business, and according to the number of link allocations between the switch groups that execute the business and the connection ports of the switch groups, a port connection relationship between the switch groups that execute the business is established, and the port connection relationship between the switch groups that execute the business is associated with the business identifier of the business.

[0199] In this way, during the execution of different services, data can be transmitted through the port connection relationship between the switch groups associated with the service identifier of the service, avoiding the problem of different tasks occupying the link. After the service execution is completed, when the port connection relationship corresponding to the service is deleted, it will not affect the execution of other services, avoiding the problem of data packet loss caused by link switching.

[0200] For example, Fig.14 As shown in the figure, two tasks are deployed on switch groups A1, A2 and A3, where the white area is task 1 and the black area is task 2. Task 1 is executed by switch groups A1, A2 and A3. The number of computing nodes occupied by switch groups A1, A2 and A3 when executing task 1 is 16, 8 and 8 respectively. Four links are allocated between switch group A1 and switch group A2 for data transmission of task 1, four links are allocated between switch group A1 and switch group A3 for data transmission of task 1, and four links are allocated between switch A2 and switch A3 for data transmission of task 1. Task 2 is executed by switch group A2 and switch A3. The number of computing nodes occupied by switch group A2 and switch group A3 when executing task 2 is 8 and 8 respectively. Therefore, 8 links are allocated between switch group A2 and switch group A3 for data transmission of task 2.

[0201] It is understandable that when both Task 1 and Task 2 have data from switch group A2 to switch group A3, Fig.14 As shown, the data of task 1 is sent to switch group A3 through the egress ports corresponding to the 4 links above switch group A2, and the data of task 2 is sent to switch group A3 through the egress ports corresponding to the 8 links below switch group A2.

[0202] In some embodiments, after obtaining the link allocation information, the controller sends the link allocation information to the optical space switching device. The optical space switching device receives the link allocation information sent by the controller, determines the optical path ports connected to each switch group through the port connection relationship indicated by the link allocation information, and forms an optical path connection relationship between the optical space switching device and each switch group based on the optical path ports connected to each switch group.

[0203] Among them, the port connection relationship includes the number of connection ports between each switch group and the connection port identification. The optical space switching device determines the optical path port connected to each switch group according to the number of connection ports between each switch group and the connection port identification. The optical space switching device forms an optical path connection relationship between the optical space switching device and each switch group based on the optical path port connected to each switch group.

[0204] Taking into account that data transmission for different tasks may need to be performed in a switch group, in order to avoid conflicts in data transmission between different tasks, in some implementations, corresponding virtual routing instances can be created for different tasks in the switch group, and the connection ports occupied by the switch group when performing different tasks are bound to the virtual routing instances of the corresponding tasks to achieve routing isolation of different tasks. A dynamic routing protocol is used to configure a routing table for the connection ports bound to the virtual routing instances of the same task, and a routing path is established between the switch groups that perform the same task to achieve data transmission of the tasks.

[0205] Among them, routing isolation means that the connection port allocated to each task in the switch group can only be used for data transmission of the task. Data generated by different tasks in the same switch group and destined for the same switch group will be sent using different links.

[0206] The virtual routing instance may be a virtual private network (VPN) instance created by the switch group through virtual routing and forwarding (VRF).

[0207] In a further implementation method, after the controller sends link allocation information to the optical space switching device in the all-optical network, it sends routing configuration instructions to the switch group that executes the service based on the link allocation information. The routing configuration instructions can be used to instruct the switch group that executes the task to perform routing configuration according to the link allocation information and establish a routing path between the switch groups that execute the task.

[0208] In yet another further implementation, the controller sends link allocation information to an optical space switching device in the all-optical network, and the optical space switching device reconstructs and configures the optical path connection between the switch groups based on the link allocation information to form an optical path connection relationship between the optical space switching device and each switch group. The optical space switching device returns optical path switching confirmation information to the controller. After receiving the optical path switching confirmation information returned by the optical space switching device, the controller sends a routing configuration instruction to the switch group that executes the service based on the link allocation information. The routing configuration instruction can be used to instruct the switch group that executes the task to perform routing configuration according to the link allocation information to establish a routing path between the switch groups that execute the task.

[0209] In some implementations, the link allocation information may include a connection port identifier of each switch group in the switch group executing the service.

[0210] In a possible implementation, in order to reduce the number of interactions between the controller and the switch group, the controller sends a routing configuration instruction to the switch group executing the service based on the connection port identifiers of each switch group in the switch group executing the service included in the link allocation information; the switch group responds to the routing configuration instruction, determines the connection port identifier of the switch group executing the service, creates a virtual routing instance in the switch group, associates the virtual routing instance with the service, and binds the virtual routing instance associated with the service to the connection port corresponding to the connection port identifier executing the service, establishes a dynamic routing protocol peer in the virtual routing path of the service, and establishes a routing path.

[0211] In a further implementation, in order to ensure the accuracy of the switch group's execution of instructions, the routing configuration instructions are divided into instance creation instructions and port binding instructions. The instance creation instruction is used to instruct the switch group to create a virtual routing instance corresponding to the service in the switch group, and the port binding instruction is used to instruct the switch group to bind the connection port of the switch group to the virtual routing instance corresponding to the service.

[0212] The controller may first send an instance creation instruction to the switch group that executes the service. The switch group responds to the instance creation instruction and creates a virtual routing instance corresponding to the service in the switch group. After completing the creation of the virtual routing instance, the switch group sends instance creation confirmation information to the controller. After receiving the instance creation confirmation information returned by the switch group that executes the service, the controller sends a port binding instruction to the switch group that executes the service based on the connection ports of each switch group in the switch group that executes the service recorded in the link allocation information. The switch group responds to the port binding instruction, determines the connection port that executes the service, binds the connection port of the switch group to the virtual routing instance corresponding to the service, establishes a dynamic routing protocol peer in the virtual routing path of the service, and establishes a routing path.

[0213] For example, Fig.15 As shown, Fig.15 This is a schematic diagram of routing isolation provided in an embodiment of the present application. Switch group S1, switch group S2 and switch group S3 execute tasks 1 and 2, create a virtual routing instance corresponding to task 1 in switch group S1, switch group S2 and switch group S3, bind connection ports 1 and 2 of switch group S1, connection ports 5 and 6 of switch group S2, and connection ports 9 and 10 of switch group S3 to the virtual routing instance corresponding to task 1; create a virtual routing instance corresponding to task 2 in switch group S1, switch group S2 and switch group S3, bind connection ports 3 and 4 of switch group S1, connection ports 7 and 8 of switch group S2, and connection ports 11 and 12 of switch group S3 to the virtual routing instance corresponding to task 2, establish BGP peers inside the virtual routing instances corresponding to tasks 1 and 2, and perform routing propagation.

[0214] like Fig.15 As shown, when switch group S1 sends the data generated by task 1 and the data generated by task 2 to switch group S2, switch group S1 sends the data generated by task 1 to connection interfaces 5 and 6 of switch group S2 through connection ports 1 and 2, and switch group S1 sends the data generated by task 2 to connection interfaces 7 and 8 of switch group S2 through connection ports 3 and 4.

[0215] Considering that when establishing routing paths between switch groups, the shortest routing paths between switch groups are usually established, that is, data is transmitted between switch groups through direct paths. When the direct links between switch groups are blocked, the data transmission of the service will be blocked. Based on this, in order to avoid the blockage of the direct links between switch groups and solve the load balancing problem, in some embodiments, when establishing routing paths between switch groups in the switch group executing the service, the shortest routing paths between switch groups and the non-shortest routing paths between switch groups are established. In the data transmission of the service, the switch group switches the currently used shortest path to a non-shortest path based on the traffic characteristics of the link when the traffic characteristics of the currently used shortest path are greater than or equal to a preset traffic characteristic threshold. This alleviates the path congestion problem caused by using only the shortest path for data transmission and further improves the link load balancing problem.

[0216] The non-shortest routing path is a data transmission between switch groups through an intermediate switch group, wherein the intermediate switch group may be a switch group that executes a service or another switch group that is in an idle state.

[0217] In a possible implementation, the switch group can create a shortest routing path and a non-shortest routing path by creating a non-shortest virtual routing instance including a first path and a second path, wherein the first path is the shortest routing path, that is, the switch group is directly connected to the remaining switch groups executing the service, and the second path is the non-shortest routing path, that is, the switch group is connected to the remaining switch groups executing the service through the intermediate switch group.

[0218] Specifically, the switch group responds to the instance creation instruction sent by the controller to create a non-shortest virtual routing instance corresponding to the service, and responds to the port binding instruction sent by the controller to bind the connection port to the non-shortest virtual routing instance to establish a routing path between each switch group.

[0219] In another possible implementation, when creating a routing virtual instance corresponding to a service, the switch group simultaneously creates a shortest virtual routing instance and a non-shortest virtual routing instance, wherein the shortest virtual routing instance includes the first path, and the non-shortest virtual routing instance includes the first path and the second path.

[0220] It can be understood that for the switch group whose connection ports are bound to the shortest virtual routing instance, data transmission is performed through the shortest routing path. For the switch group whose connection ports are bound to non-shortest virtual routing instances, when the shortest routing path is unobstructed, data transmission is performed through the shortest routing path. When the shortest routing path is blocked, data transmission is performed through the non-shortest routing path.

[0221] In a possible implementation manner of the above embodiment, the switch group may bind all connection ports executing the service to the non-shortest virtual routing instance corresponding to the service.

[0222] In another possible implementation of the above embodiment, the switch group can bind the connection port executing the service with the virtual routing instance to be bound according to the virtual routing instance to be bound specified in the port binding instruction. The virtual routing instance to be bound includes the shortest virtual routing instance and the non-shortest virtual routing instance.

[0223] Specifically, when the controller sends a port binding instruction to the switch group executing the service, it specifies in the port binding instruction which virtual routing instance the connection port in the switch group needs to be bound to. The switch group responds to the port binding instruction, determines the connection port executing the service and the virtual routing instance to be bound to the connection port executing the service, and binds the virtual routing instance to be bound to the connection port executing the service with the connection port executing the service.

[0224] For example, the port binding instruction can indicate which virtual routing instance the connection port in the switch group needs to be bound to through the instance identifier. Specifically, the port binding instruction carries the port identifier and the instance identifier of the connection port executing the service in the switch group. After responding to the port binding instruction, the switch group binds the virtual routing instance corresponding to the instance identifier to the connection port corresponding to the port identifier. Among them, the instance identifier is used to instruct the switch group to bind the connection port to the shortest virtual routing instance or the non-shortest virtual routing instance. In some embodiments, the instance identifier can be a set of strings or identifiers pre-agreed by the switch group and the controller.

[0225] In another possible implementation of the above embodiment, the switch group may bind the virtual routing instance corresponding to the port type and the connection port corresponding to the port identifier according to the port identifier and the port type of the connection port executing the service specified in the port binding instruction. The port type is used to characterize whether the connection port is connected to a computing node, and includes a first type and a second type, wherein a port of the first type characterizes that the connection port is a port connected to a computing node, and a port of the second type characterizes that the connection port is a port connected to another switch group.

[0226] It is understandable that the port binding instruction carries the port identifier and port type of the connection port executing the service in the switch group.

[0227] For example, in order to reduce the amount of calculation of the switch group and reduce the network resource occupation caused by transmitting data on non-shortest routing paths, the connection ports in the switch group connected to the computing nodes can be bound to the non-shortest virtual routing instance, and the connection ports in the switch group connected to other switch groups can be bound to the shortest virtual routing instance.

[0228] Specifically, the controller determines the connection ports in the switch group and the port types of the connection ports according to the link allocation information between the switch groups and the node mapping relationship between the computing nodes executing the services, and sends a port binding instruction to the switch group based on the port identifiers and port types of the connection ports in the switch group. The switch group responds to the port binding instruction, binds the connection ports of the first type to the non-shortest virtual routing instance, and binds the connection ports of the second type to the shortest virtual routing instance.

[0229] For example, Fig.16As shown, in switch group S1, switch group S2 and switch group S3, the connection ports connected to the computing nodes in switch group S1 are 1 to 4, the connection ports connected to the computing nodes in switch group S2 are 5 to 6, and the connection ports connected to the computing nodes in switch group S3 are 9 to 12. Connection ports 5 and 6 of switch group S1 are connected to connection ports 7 and 8 of switch group S2, connection ports 7 and 8 of switch group S1 are connected to connection ports 7 and 8 of switch group S3, and connection ports 5 and 6 of switch group S3 are connected to connection ports 9 and 10 of switch group S2.

[0230] Switch group S1, switch group S2 and switch group S3 respectively respond to the instance creation instruction issued by the controller, and create virtual routing instances: mix_vrf and min_vrf in switch group S1, switch group S2 and switch group S3, where mix_vrf is a non-shortest virtual routing instance and min_vrf is the shortest virtual routing instance.

[0231] Switch group S1, switch group S2 and switch group S3 respond to the port binding instruction issued by the controller respectively. Switch group S1 binds connection ports 1 to 4 in switch group S1 to the non-shortest virtual routing instance mix_vrf, and binds connection ports 7, 8, 5, and 6 to the shortest virtual routing instance min_vrf; switch group S2 binds connection ports 5 to 8 in switch group S2 to the non-shortest virtual routing instance mix_vrf, and binds connection ports 7, 8, 9, and 10 to the shortest virtual routing instance min_vrf; switch group S3 binds connection ports 9 to 12 in switch group S3 to the non-shortest virtual routing instance mix_vrf, and binds connection ports 7, 8, 5, and 6 to the shortest virtual routing instance min_vrf. Fig.16 shown.

[0232] When the data generated by the service goes from switch group S1 to switch group S2, the data generated by the service in the computing node enters switch group S1 from connection ports 1 to 4 of switch group S1. Since connection ports 1 to 4 of switch group S1 are bound to non-shortest virtual routing instances, the data can be transmitted to switch group S2 through connection ports 5 and 6 of switch group S1, or to switch group S3 through connection ports 7 and 8 of switch group S1. Since the connection port in switch group S3 connected to switch group S1 is bound to the shortest virtual routing instance, switch group S3 transmits the data to switch group S2 through the shortest routing path. That is, when the data generated by the service goes from switch group S1 to switch group S2, the data can be transmitted through the shortest routing path S1-S2, or through the non-shortest routing path S1-S3-S2. In some embodiments, after establishing a routing path, the switch group executing the service sends routing path confirmation information to the controller. After receiving the routing path confirmation information returned by the switch group executing the service, the controller triggers the service execution operation, executes the service through the computing node assigned to the service, and realizes the transmission of the data generated by the service through the routing path between the switch groups executing the service.

[0233] In a possible implementation, for a switch group whose connection ports are bound to the shortest virtual routing instance, when executing the transmission of data generated by the service, the data is transmitted through the shortest routing path.

[0234] In one possible implementation, for a switch group whose connection port is bound to a non-shortest virtual routing instance, when executing the transmission of data generated by the service, data is transmitted through the shortest routing path, and the traffic characteristics of the shortest routing path are monitored. When the traffic characteristics on the shortest routing path are greater than a preset traffic characteristic threshold, that is, when the shortest routing path is congested, the shortest routing path is switched to a non-shortest routing path, and data is transmitted through the non-shortest routing path.

[0235] When the switch group transmits data generated by the service in the form of data packets, when it is detected that the shortest routing path is congested, the data packets to be sent are transmitted through a non-shortest routing path.

[0236] Among them, when the switch group transmits data generated by the business in the form of data segments, when the shortest routing path is detected to be congested, part of the data segments in the data to be sent are transmitted through a non-shortest routing path, and the remaining data segments are transmitted through the shortest routing path.

[0237] Based on this implementation, the switch group creates a non-shortest virtual routing instance that includes a non-shortest routing path and a shortest routing path. When the shortest routing path is congested during data transmission, data is transmitted through the non-shortest routing path to alleviate the load balancing problem.

[0238] In order to ensure that the reconstruction and disconnection of the connection links between the switch groups executing the services in the all-optical network will not affect the execution of other services, and at the same time, improve the resource utilization of the all-optical network. In some embodiments, after the controller obtains the link allocation information of the switch group executing the service, the link allocation information is associated with the service identifier of the service. In this way, after the service is executed, the computing node executing the service (i.e., the service that has been executed is completed) is released, and the port connection relationship corresponding to the service identifier of the service in the optical space switching device is deleted based on the link allocation information associated with the service identifier of the service, and the virtual routing instance corresponding to the service in the switch group is deleted.

[0239] Specifically, the method may also include: the scheduler monitors the execution status of the service in the all-optical network, and when it is monitored that the service execution is completed, the scheduler takes the executed service as the target service, and sends a link release request to the controller based on the service identifier of the target service. The controller responds to the link release request, obtains the target service identifier of the target service carried in the link release request, determines the target link allocation information associated with the target service identifier, sends a link deletion request to the optical space switching device based on the target link allocation information, and sends a routing path deletion instruction to the target switch group executing the target service. The optical space switching device receives and responds to the link deletion request, and deletes the port connection relationship indicated by the target link allocation information. The target switch group executing the target service receives and responds to the routing path deletion instruction, and deletes the target virtual routing instance corresponding to the target service in the target switch group.

[0240] Among them, the execution status includes waiting for execution, executing, and execution completed. When the scheduler detects that the execution status of the service is service execution completed, the service with the execution status of service execution completed is regarded as the execution completed service, and a link release request is sent to the controller based on the service identifier of the execution completed service.

[0241] Among them, the link release request can be used to request the release of the link that transmits the data of the target service, the link deletion request can be used to instruct the optical space switching device to delete the port connection relationship indicated by the target link allocation information, and the routing path deletion instruction can be used to instruct the deletion of the target virtual routing instance corresponding to the target service in the target switch group.

[0242] In this way, after the service execution is completed, the link of the service in the all-optical network can be cut off. Since the links of different services in the switch group are different, when the link of the service in the all-optical network is cut off, it will not affect the data transmission of other running services. At the same time, the link of the completed service is released to improve resource utilization.

[0243] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of interaction between each node. It is understandable that each node, such as a controller, an optical space switching device, etc., includes a hardware structure and / or software module corresponding to each function in order to realize the above functions. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0244] The embodiment of the present application can group the functional modules of the controller, optical space switching device, etc. according to the above method example. For example, each functional module can be grouped according to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the grouping of modules in the embodiment of the present application is schematic and is only a logical functional grouping. There may be other grouping methods in actual implementation.

[0245] In order to better implement the link reconstruction method provided in the embodiment of the present application, based on the above link reconstruction method embodiment, the embodiment of the present application provides a link reconstruction device provided in the controller, such as Fig.17 As shown, Fig.17 1 is a schematic diagram of the structure of a link reconstruction device 17 provided in a controller according to an embodiment of the present application. The schematic diagram of the structure of the link reconstruction device 17 provided in the controller includes:

[0246] The acquisition module 171 is used to acquire resource information of the service; the resource information is used to indicate each switch group executing the service in the all-optical network and the computing power parameters of each switch group; the computing power parameters are used to indicate the computing power of the switch group when executing the service;

[0247] The processing module 172 is used to determine link allocation information according to the computing power parameters of each switch group; the link allocation information is used to indicate the port connection relationship between the switch groups executing the service;

[0248] The sending module 173 is used to send link allocation information to the optical space switching device in the all-optical network.

[0249] Among them, the acquisition module 171, the processing module 172 and the sending module 173 can all be implemented by software or by hardware. Exemplarily, the implementation of the processing module 172 is introduced below by taking the processing module 172 as an example. Similarly, the acquisition module 171 and the sending module 173 can refer to the implementation of the processing module 172.

[0250] As an example of a software functional unit, the processing module 172 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance may be one or more. For example, the processing module 172 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with similar geographical locations. Generally, a region may include multiple AZs.

[0251] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0252] As an example of a hardware functional unit, the processing module 172 may include at least one computing device, such as a server, etc. Alternatively, the processing module 172 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0253] The multiple computing devices included in the processing module 172 can be distributed in the same region or in different regions. The multiple computing devices included in the processing module 172 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the processing module 172 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0254] It should be noted that the naming and grouping of devices and modules in the embodiments of the present application are schematic and are only a logical function grouping. There may be other grouping methods in actual implementation. For example, the link reconstruction device 17 may also be named a topology engineering module or other names without limitation. In addition, the link reconstruction device 17 may not only be divided into an acquisition module 171, a processing module 172, and a sending module 173 according to its execution function, but also may be divided into a topology calculation module 211 and a routing calculation module 212. Specifically, the other division example may refer to the following Fig. 20 As described in the corresponding embodiments.

[0255] It should be noted that, in other embodiments, the processing module 172 can be used to execute any step in the link reconstruction method applicable to the controller, the acquisition module 171 can be used to execute any step in the link reconstruction method applicable to the controller, and the sending module 173 can be used to execute any step in the link reconstruction method applicable to the controller. The steps that the acquisition module 171, the processing module 172, and the sending module 173 are responsible for implementing can be specified as needed, and the acquisition module 171, the processing module 172, and the sending module 173 respectively implement different steps in the link reconstruction method applicable to the controller to implement all functions of the link reconstruction device set in the controller.

[0256] In order to better implement the link reconstruction method provided in the embodiment of the present application, based on the above link reconstruction method embodiment, the embodiment of the present application provides a link reconstruction device 18 provided in an optical space switching device, such as Fig.18 As shown, Fig.18 1 is a schematic diagram of the structure of a link reconstruction device 18 provided in an optical space switching device according to an embodiment of the present application. The schematic diagram of the structure of the link reconstruction device 18 provided in the optical space switching device shown includes:

[0257] The receiving module 181 is used to receive link allocation information from the controller; the link allocation information is resource information for the controller to obtain services, and the optical path connection between the switch groups is reconfigured according to the link allocation information to form an optical path connection relationship between the optical space switching device and each switch group; the resource information is used to indicate each switch group that executes services in the all-optical network and the computing power parameters of each switch group, and the link allocation information is determined according to the computing power parameters of each switch group, and is used to indicate the port connection relationship between the switch groups that execute services;

[0258] The link processing module 182 is used to reconfigure the optical path connection between the switch groups according to the link allocation information, so as to form an optical path connection relationship between the optical space switching device and each switch group.

[0259] The link reconstruction device provided in the optical space switching device according to the embodiment of the present application establishes an optical path connection relationship between the optical space switching device and each switch group based on the link allocation information sent by the controller, and deletes the port connection relationship indicated by the target link allocation information carried by the link deletion request in response to the link deletion request sent by the controller, and realizes the reconstruction of the link in the all-optical network through interaction with the controller.

[0260] like Fig.19 As shown, an embodiment of the present application also provides a structural diagram of a cluster that applies the link reconstruction method, wherein the cluster includes an all-optical network 19, a controller 422, and a scheduler 421, and the all-optical network 19 includes an optical space switching device 4231, a switch group 4232, and a computing node 4233.

[0261] The controller 422 is connected to the optical space switching device 4231, the switch group 4232, and the scheduler 421 respectively. Fig.17The provided link reconstruction device 17, the controller 422 can execute the link reconstruction method executed by the above-mentioned link reconstruction device 17, for example: the controller 422 can be used to obtain resource information of the service, determine the link allocation information according to the computing power parameters of each switch group 4232, send the link allocation information to the optical space switching device 4231 in the all-optical network 19, and send the routing configuration instruction to the switch group 4232 based on the link allocation information. For the specific execution process, reference can be made to the embodiment of the above-mentioned link reconstruction method, and the embodiment of the present application will not be repeated here.

[0262] The downlink port of the optical space switching device 4231 is connected to the uplink port of the switch group 4232. Fig.18 The provided link reconstruction device 18 is used to receive link allocation information from the controller 422, and reconstruct the optical path connection between the switch groups 4232 according to the link allocation information to form an optical path connection relationship between the optical space switching device 4231 and each switch group 4232. The specific execution process can refer to the above-mentioned link reconstruction method embodiment, and the embodiment of the present application will not be repeated here.

[0263] The downlink port of the switch group 4232 is connected to a computing node 4233. The switch groups 4232 are interconnected through the optical space switching device 4231, and the computing nodes 4233 communicate through the switch groups 4232. The switch group 4232 is used to receive and respond to the routing configuration instructions sent by the controller 422 to establish a routing path.

[0264] The scheduler 421 is coupled to the controller 422, or the scheduler 421 and the controller 422 are integrated into one device, or the scheduler 421 is set as a processing unit in the controller 422. The scheduler 421 is used for job scheduling and resource allocation of services, forms resource information of services, and sends the resource information to the controller 422.

[0265] In some embodiments, the controller 422 is a device for controlling the optical space switching device 4231 and the switch group 4232 in the all-optical network 19. The controller 422 is also an abstract and general term and does not represent a real hardware device. In actual implementation, the controller 422 may include only one hardware device, which may mainly include a general-purpose processor or a field programmable gate array (FPGA). At the same time, it is also necessary to include a corresponding interface to connect with the optical space switching device 4231 and the switch group 4232 to control the optical space switching device 4231 and the switch group 4232.

[0266] In one example, the controller may include: Fig.17 The link reconstruction device 17 shown may include an acquisition module 171 , a processing module 172 and a sending module 173 .

[0267] In another example, for example, Fig. 20 As shown, the controller 422 is provided with a topology engineering module 21, which is connected to the scheduler 421, receives the resource information of the service sent by the scheduler 421, and performs link reconstruction according to the link reconstruction method applicable to the controller 422 based on the resource information of the service.

[0268] In an optional embodiment, a topology calculation module 211 and a routing calculation module 212 are provided in the topology engineering module 21, wherein the topology calculation module 211 is used to receive resource information of the service sent by the scheduler 421, obtain link allocation information and node mapping relationship based on the resource information of the service according to the above-mentioned link reconstruction method applicable to the controller 422, send the link allocation information to the optical space switching device 4231, and store the node mapping relationship; the routing calculation module 212 is used to send routing configuration instructions to the switch group 4232 according to the link allocation information obtained by the topology engineering module 21.

[0269] The topology engineering module 21 and the routing calculation module 212 may be software function modules including codes running on a computing instance, or may be hardware function units including at least one computing device.

[0270] In some implementations, the controller 422 may include two or more physical hardware devices, which are respectively connected to the optical space switching device 4231 and the switch group 4232 to control the optical space switching device 4231 and the switch group 4232. These hardware devices can communicate with each other and work together to achieve control of the entire all-optical network 19.

[0271] like Fig.21 As shown, the controller 422 may include a processor 2202 and a channel interface 2201 .

[0272] In a possible implementation, the controller 422 may further include a memory 2203 for storing a program. Specifically, the program may include a program code, and the program code includes a computer operation instruction. The memory 2203 may include a read-only memory and a random access memory, and provide instructions and data to the processor 2202. The memory 2203 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage.

[0273] The channel interface 2201, the processor 2202 and the memory 2203 are interconnected through the bus 2204 system. The bus 2204 can be an industry standard architecture bus (ISA), a peripheral component interconnect bus (PCI) or an extended industry standard architecture bus (EISA). The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.21 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0274] The link reconstruction method disclosed in the above method embodiment may be applied to the processor 2202, or implemented by the processor 2202. The processor 2202 may be an integrated circuit chip having a signal processing capability.

[0275] In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 2202. The above processor 2202 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete electron tubes or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or the hardware and software modules in the decoding processor are combined and executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 2203, and the processor 2202 reads the information in the memory 2203 and completes the steps of the above method in combination with its hardware.

[0276] In a possible implementation, the processor 2202 may also be used to execute a link reconstruction method. For a specific implementation, reference may be made to the embodiments provided in the above-mentioned link reconstruction method, and the embodiments of the present application will not be described in detail herein.

[0277] In the embodiment of the present application, the chip system may be composed of a chip, or may include a chip and other discrete devices.

[0278] The embodiment of the present application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be completed by a computer program to instruct the relevant hardware, and the program can be stored in the above computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a terminal of any of the above embodiments, such as: an internal storage unit including a data transmission end and / or a data receiving end, such as a hard disk or memory of a terminal. The above computer-readable storage medium can also be an external storage device of the above terminal, such as a plug-in hard disk equipped on the above terminal, a smart memory card (smart med ia card, SMC), a secure digital (secure digitala l, SD) card, a flash card (flash card), etc. Further, the above computer-readable storage medium can also include both the internal storage unit of the above terminal and an external storage device. The above computer-readable storage medium is used to store the above computer program and other programs and data required by the above terminal. The above computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.

[0279] It should be understood that the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution of this application are in compliance with relevant laws and regulations and do not violate public order and good morals. For example, in the technical solution of this application, the processing of user personal information is carried out with the authorization of the user, and the same description is not repeated here.

[0280] It should be noted that the terms "first" and "second" in the specification, claims and drawings of the present application are used to distinguish different objects rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices.

[0281] It should be understood that in the present application, "at least one (item)" means one or more, "more than one" means two or more, "at least two (items)" means two or three and more than three, and "and / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0282] It should be understood that in the embodiments of the present application, "B corresponding to A" means that B is associated with A. For example, B can be determined based on A. It should also be understood that determining B based on A does not mean determining B only based on A, but B can also be determined based on A and / or other information. In addition, the "connection" that appears in the embodiments of the present application refers to various connection methods such as direct connection or indirect connection to achieve communication between devices, and the embodiments of the present application do not impose any limitation on this.

[0283] Unless otherwise specified, the "transmission" (transmit / transmission) appearing in the embodiments of the present application refers to bidirectional transmission, including sending and / or receiving actions. Specifically, the "transmission" in the embodiments of the present application includes the sending of data, the receiving of data, or the sending of data and the receiving of data. In other words, the data transmission here includes uplink and / or downlink data transmission. Data may include channels and / or signals, uplink data transmission is uplink channel and / or uplink signal transmission, and downlink data transmission is downlink channel and / or downlink signal transmission. The "network" and "system" appearing in the embodiments of the present application express the same concept, and an all-optical network is an all-optical system.

[0284] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the grouping of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be grouped into different functional modules to complete all or part of the functions described above.

[0285] In the several embodiments provided in the present application, it should be understood that the disclosed communication devices and methods can be implemented in other ways. For example, the communication device embodiments described above are only schematic. For example, the grouping of the modules or units is only a logical function grouping. There may be other grouping methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0286] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0287] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0288] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device, such as a single-chip microcomputer, a chip, etc., or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media for storing program codes such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks or optical disks.

[0289] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A link reconstruction method, characterized in that: Applied in an all-optical network, the method is executed by a controller, and the method includes: Acquire resource information of the service; the resource information is used to indicate each switch group executing the service in the all-optical network and a computing power parameter of each switch group; the computing power parameter is used to indicate the computing power of the switch group when executing the service; Determine link allocation information according to the computing power parameters of each of the switch groups; the link allocation information is used to indicate the port connection relationship between the switch groups executing the service; The link allocation information is sent to an optical space switching device in the all-optical network.

2. The method according to claim 1, characterized in that The port connection relationship is used to indicate the number of connection ports and connection port identifiers between each of the switch groups; The port connection relationship is established based on the number of link allocations between each of the switch groups and the port information of each of the switch groups; the number of connection ports between each of the switch groups is the same as the number of link allocations between each of the switch groups; wherein the number of link allocations between each of the switch groups is allocated in proportion to the computing power parameters of each of the switch groups.

3. The method according to claim 2, characterized in that The computing power parameter of the switch group includes the number of computing nodes occupied when executing the service; The number of links allocated between the switch groups is allocated according to the ratio of the number of computing nodes occupied by each switch group when executing the service.

4. The method according to claim 3, characterized in that The method further comprises: Selecting the first switch group from the switch groups executing the service in descending order according to the number of computing nodes occupied by each switch group when executing the service; Determine the link allocation ratio between the switch groups according to the ratio between the number of computing nodes occupied by the first switch group when executing the service and the number of computing nodes occupied by the remaining switch groups when executing the service; the link allocation ratio is used to indicate the ratio of the number of computing nodes between the switch group and the remaining switch groups other than the switch group in the switch groups executing the service; The link allocation quantity between the switch groups is obtained according to the link allocation ratio between the switch groups and the maximum link allocation quantity of the switch groups; the maximum link allocation quantity of the switch group is equal to the number of computing nodes occupied by the switch group when executing the service.

5. The method according to claim 4, characterized in that The obtaining the link allocation quantity between the switch groups according to the link allocation ratio between the switch groups and the maximum link allocation quantity of the switch groups comprises: According to the link allocation ratio between the switch group and each second switch group, the maximum link allocation quantity of the switch group is allocated to each second switch group to obtain the link allocation quantity between the switch group and each second switch group; the second switch group is the switch group other than the switch group in the switch group executing the service; the sum of the link allocation quantity between the switch group and each second switch group is less than or equal to the maximum link allocation quantity of the switch group.

6. The method according to claim 5, characterized in that The step of allocating the maximum link allocation quantity of the switch group to each of the second switch groups to obtain the link allocation quantity between the switch group and each of the second switch groups includes: Allocate the maximum link allocation quantity of the switch group to each second switch group according to the link allocation ratio between the switch group and each second switch group, and determine the initial link allocation quantity between the switch group and each second switch group; If the sum of the initial link allocation quantities between the switch group and each of the second switch groups is equal to the maximum link allocation quantity of the switch group, the initial link allocation quantity between the switch group and each of the second switch groups is determined as the link allocation quantity between the switch group and each of the second switch groups; If the sum of the initial link allocation quantities between the switch group and each of the second switch groups is less than the maximum link allocation quantity of the switch group, the initial link allocation quantity between the switch group and the third switch group is adjusted to obtain the link allocation quantity between the switch group and each of the second switch groups; the third switch group is included in the second switch group whose link allocation quantity is less than the maximum link allocation quantity.

7. The method according to claim 6, characterized in that The third switch group is selected from the second switch groups whose link allocation number is less than the maximum link allocation number based on the traffic characteristics between the switch group and each of the second switch groups; the traffic characteristics are used to indicate the data transmission volume between the switch groups when executing the service.

8. The method according to claim 1, characterized in that After sending the link allocation information to the optical space switching device in the all-optical network, the method further includes: A routing configuration instruction is sent to each of the switch groups according to the link allocation information; the routing configuration instruction is used to instruct each of the switch groups to perform routing configuration according to the link allocation information to establish a routing path between each of the switch groups.

9. The method according to claim 8, characterized in that The link allocation information is also used to indicate the connection port of each switch group in the switch group executing the service; The configuration instruction includes an instance creation instruction and a port binding instruction; the instance creation instruction is used to instruct the switch group to create the virtual routing instance corresponding to the service in the switch group; the port binding instruction is used to instruct the switch group to bind the connection port of the switch group with the virtual routing instance corresponding to the service; The instance creation instruction is sent to the switch group executing the service based on the switch group executing the service indicated by the link allocation information; The port binding instruction is sent to the switch group executing the service based on the identification information of the connection port of each switch group after receiving the creation confirmation information returned by the switch group executing the service based on the instance creation instruction.

10. The method according to claim 9, characterized in that The virtual routing instances include the shortest virtual routing instances and non-shortest virtual routing instances.

11. The method according to claim 9, characterized in that The service identifier of the service is associated with the link allocation information; the method further includes: In response to a link release request, obtaining a target service identifier of a target service carried in the link release request; the link release request is used to request the release of a link for transmitting data of the target service; Determining target link allocation information associated with the target service identifier; Sending a link deletion request to the optical space switching device based on the target link allocation information to delete the port connection relationship corresponding to the target link allocation information in the all-optical network, wherein the link deletion request is used to instruct the optical space switching device to delete the port connection relationship indicated by the target link allocation information; A routing path deletion instruction is sent to a target switch group of the target service, where the routing path deletion instruction is used to instruct to delete a target virtual routing instance corresponding to the target service in the target switch group.

12. The method according to any one of claims 1 to 11, characterized in that: The resource information of the business is obtained, including: In response to the link application request, resource information of the service sent by the scheduler 421 is obtained; the resource information is obtained by the scheduler by responding to the service application request, performing task scheduling based on the service resources carried in the service application request and the currently idle computing nodes in the all-optical network.

13. A link reconstruction method, characterized in that: An optical space switching device applied to an all-optical network, the method comprising: The optical space switching device receives link allocation information from the controller; the link allocation information is resource information for the controller to obtain the service; the resource information is used to indicate each switch group executing the service in the all-optical network and the computing power parameters of each switch group, and the link allocation information is determined based on the computing power parameters of each switch group and is used to indicate the port connection relationship between the switch groups executing the service; The optical space switching device reconfigures the optical path connections between the switch groups according to the link allocation information, thereby forming an optical path connection relationship between the optical space switching device and each of the switch groups.

14. The method according to claim 13, characterized in that The port connection relationship is used to indicate the number of connection ports and connection port identifiers between the switch groups; the optical space switching device reconfigures the optical path connection between the switch groups according to the link allocation information to form an optical path connection relationship between the optical space switching device and each of the switch groups, including: The optical space switching device determines the optical path port connected to each of the switch groups according to the number of connection ports between the switch groups and the connection port identifiers; The optical space switching device configures the optical path connections between the switch groups based on the optical path ports connected to the switch groups, thereby forming an optical path connection relationship between the optical space switching device and the switch groups.

15. The method according to claim 13 or 14, characterized in that The method further comprises: The optical space switching device responds to the link deletion request sent by the controller, and based on the target link allocation information carried by the link deletion request, deletes the port connection relationship indicated by the target link allocation information; the link deletion request is used to instruct the optical space switching device to delete the port connection relationship indicated by the target link allocation information.

16. The method according to claim 15, characterized in that The optical space switching device deletes the port connection relationship indicated by the target link allocation information, including: disconnecting the optical path port connected to each target switch group, and deleting the port connection relationship indicated by the target link allocation information.

17. A link reconstruction method, characterized in that: Applied to a cluster including an all-optical network and a controller, the all-optical network including an optical space switching device group and a switch group, the method includes: The controller obtains resource information of the service, determines link allocation information according to the computing power parameters of each switch group, and sends the link allocation information to the optical space switching device in the all-optical network; wherein the resource information includes each switch group executing the service in the all-optical network and the computing power parameters of each switch group, and the computing power parameters are used to indicate the computing power of the switch group when executing the service; the link allocation information is used to indicate the port connection relationship between the switch groups executing the service; The optical space switching device receives the link allocation information from the controller, and reconfigures the optical path connections between the switch groups according to the link allocation information to form an optical path connection relationship between the optical space switching device and each of the switch groups.

18. The method according to claim 17, characterized in that The cluster further includes a scheduler, and the method further includes: The scheduler responds to the service application request, performs task scheduling based on the service resources carried in the service application request and currently idle computing nodes in the all-optical network, obtains service resource information, and sends the service resource information to the controller.

19. The method according to claim 17 or 18, characterized in that The method further comprises: The controller sends a routing configuration instruction to each of the switch groups according to the link allocation information; The switch group responds to the routing configuration instruction, performs routing configuration according to the link allocation information, and establishes routing paths between the switch groups.

20. The method according to claim 19, characterized in that The routing configuration instruction includes an instance creation instruction and a port binding instruction, and the method further includes: The switch group creates a virtual routing instance corresponding to the service in response to the instance creation instruction sent by the controller, and returns instance creation confirmation information to the controller; The switch group determines the connection port to be bound in response to the port binding instruction sent by the controller, binds the connection port to the virtual routing instance, and establishes a routing path between each of the switch groups.

21. The method according to claim 20, characterized in that The virtual routing instance includes a shortest virtual routing instance and a non-shortest virtual routing instance; the connection port includes a first type of connection port and a second type of connection port, the first type of connection port represents that the connection port is a port connected to the computing node, and the second type of connection port represents that the connection port is a port connected to other switch groups; The first type of connection port is bound to the non-shortest virtual routing instance, and the second type of connection port is bound to the shortest virtual routing instance.

22. The method according to claim 20, characterized in that After establishing the routing paths between the switch groups, the method further includes: The switch group receives data corresponding to the service, and forwards the data to other switch groups executing the service in the all-optical network through the virtual routing instance corresponding to the service.

23. The method according to claim 22, characterized in that The virtual routing instance includes a shortest virtual routing instance and a non-shortest virtual routing instance, the shortest virtual routing instance includes a first path, the non-shortest virtual routing instance includes a first path and a second path, the first path represents that the switch group is directly connected to the remaining switch groups executing the service, and the second path represents that the switch group is communicated and connected with the remaining switch groups executing the service through the intermediate switch group; The shortest virtual routing instance is used to forward the data to the remaining switch groups executing the service in the all-optical network through the first path; The non-shortest virtual routing instance is used to forward the data to the remaining switch groups that execute the service in the all-optical network through the second path when the traffic characteristic of the first path is greater than a preset traffic characteristic threshold; when the traffic characteristic of the first path is less than or equal to the preset traffic characteristic threshold, forward the data to the remaining switch groups that execute the service in the all-optical network through the first path; the traffic characteristic is used to indicate the quantity transmission volume between the switch groups when executing the service.

24. A link reconstruction device, characterized in that: The device comprises: An acquisition module, used to acquire resource information of a service; the resource information is used to indicate each switch group executing the service in the all-optical network and a computing power parameter of each switch group; the computing power parameter is used to indicate the computing power of the switch group when executing the service; A processing module, configured to determine link allocation information according to the computing power parameters of each of the switch groups; the link allocation information is used to indicate the port connection relationship between the switch groups executing the service; The sending module is used to send the link allocation information to the optical space switching device in the all-optical network.

25. A controller, characterized in that: The controller includes a processing module for executing the link reconstruction method described in any one of claims 1 to 12, or the link reconstruction device described in claim 24 is deployed in the controller.

26. A link reconstruction device, characterized in that: The device comprises: A receiving module, configured to receive link allocation information from a controller; the link allocation information is resource information for the controller to obtain services, and according to the link allocation information, the optical path connection between the switch groups is reconfigured to form an optical path connection relationship between the optical space switching device and each of the switch groups; the resource information is used to indicate each switch group executing the service in the all-optical network and the computing power parameters of each of the switch groups, and the link allocation information is determined based on the computing power parameters of each of the switch groups, and is used to indicate the port connection relationship between the switch groups executing the service; The link processing module is used to reconfigure the optical path connection between the switch groups according to the link allocation information to form an optical path connection relationship between the optical space switching device and each of the switch groups.

27. An optical space switching device, characterized in that: The optical space switching device is used to execute the link reconstruction method described in any one of claims 13 to 16, or the link reconstruction device described in claim 26 is deployed in the optical space switching device.

28. A cluster, characterized in that: The cluster includes an all-optical network, a controller and a scheduler, and the all-optical network includes an optical space switching device and a switch group; the cluster is used to execute the method according to any one of claims 17 to 23.