Data transmission method and cluster system
By finely isolating faults and dynamically adjusting redundant resources in the cluster system, the contradiction between system capacity and redundancy in the multi-frame cluster architecture is solved, the system's operating efficiency and availability are improved, and the risk of business interruption is reduced.
Patent Information
- Application Number
- CN202410200235.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-22
- Publication Date
- 2025-08-22
AI Technical Summary
There is a contradiction between improving system capacity and redundancy in multi-frame cluster architecture equipment. The existing technology cannot effectively allocate redundant resources without increasing equipment volume and reducing system capacity, resulting in high risk of business interruption and low system availability.
By finely isolating faults in the cluster system, the service data associated with the fault board is controlled not to go to the central switching board connected to the fault interface board, and the logical port switching and bandwidth monitoring of the cross board are used to dynamically adjust redundant resources to avoid interruptions in the entire frame of business.
It improves the utilization rate of redundant resources, improves the operating efficiency and availability of cluster systems, reduces the impact of failures, and realizes the ability to flexibly adjust redundant resources based on the actual bandwidth of the service.
Smart Images

Figure CN120528512A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of optical communication technology, and more specifically, to a data transmission method and a cluster system. Background Art
[0002] With the development of informatization and cloud computing, demand for dedicated lines and video services is growing rapidly. These services are characterized by low bandwidth and high volume, requiring simple, fast, and flexible bandwidth adjustments. Optical transport networks (OTNs) are widely deployed in backbone lines, metropolitan cores, and metropolitan edge networks, offering inherent advantages of high quality, large capacity, and wide coverage.
[0003] However, the multi-chassis cluster architecture creates a conflict between increasing the capacity and redundancy of the overall communication system. System capacity is constrained by process technology and integration, making it impossible to simultaneously increase both capacity and redundancy. Furthermore, the availability of fault redundancy protection solutions is low.
[0004] Therefore, how to reasonably allocate system redundant resources, reduce the risk of business interruption, and improve system availability without increasing equipment size or reducing system capacity is one of the research hotspots in this field. Summary of the Invention
[0005] The present application provides a method and cluster system for transmitting data. In response to a failure of any interface board, the services associated with the faulty board are controlled not to be sent to the faulty interface board, and the transmission of other unrelated services is not affected, without the need to isolate the entire switching plane. By fine-tuning fault isolation, the utilization rate of redundant resources is improved. At the same time, the method and system provided by the present application also have the ability to control the issuance and recovery of services, and flexibly adjust the number of redundant resources according to the actual bandwidth of the service; when the number of faulty boards exceeds the protection capacity, the service bandwidth exceeding the protection capacity is actively recovered, thereby increasing the number of redundant resources, avoiding interruption of services in the entire frame, and reducing the impact of the failure.
[0006] In a first aspect, a method for transmitting data is provided. The method is performed in a cluster system, wherein the cluster system includes N service boards, N interface board sets, and M central switching boards, where N is greater than or equal to 3 and M is greater than or equal to 2. Each interface board set includes at least one interface board, and there is a one-to-one correspondence between the N service boards and the N interface board sets. Each service board sends or receives service data through an interface board in its corresponding interface board set. Each central switching board is connected to at least three interface boards, each of which belongs to different interface board sets. In response to a failure of a first interface board connected to a first central switching board among the M central switching boards, the method controls the first service board to transmit first service data to a second service board based on a first policy, and controls the third service board to transmit second service data based on a second policy. The first interface board belongs to the interface board set corresponding to the first service board or the second service board. The first policy is used to indicate a first target central switching board set for the first service data, and the second policy is used to indicate a second target central switching board set for the second service data. The first target central switching board set includes central switching boards that are connected to at least one interface board other than the first interface board in the interface board set to which the first interface board belongs. The second target central switching board set includes the first central switching board.
[0007] Based on the above solution, in response to any interface board failure, the present application prevents service data associated with the failed board from being sent to the central switching board connected to the failed interface board. However, service data transmitted between the central switching board and other normal interface boards remains unaffected, eliminating the need to isolate the entire switching plane. This refined fault isolation improves the utilization of redundant resources and enhances the overall cluster system efficiency.
[0008] It should be understood that the central switching board is connected to at least two interface board sets, and the central switching board is connected to at least one interface board in the interface board set, forming a switching plane. There is no service data exchange between the switching planes, and there is no service data exchange between the central switching boards. These are virtual parallel switching planes.
[0009] It should be understood that due to load balancing, in response to no faults in the interface boards in the system, the first service board transmits the first service data between the second service board and the central switching board connected to all interface boards in the first interface board set. The first interface board set corresponds to the first service board and includes at least one interface board connected to the second service board. The first interface board set includes interface boards connected to the first central switching board set, which includes central switching boards involved in the transmission of the first service data.
[0010] It should be understood that the second service data includes service data transmitted between the third service board and the remaining service boards except the service board connected to the first interface board.
[0011] It should be understood that the first interface board includes one or more interface boards, and this application does not impose any special limitation on this.
[0012] In a specific implementation, the first interface board belongs to an interface board set corresponding to the first service board, and the second service data includes service data transmitted between the third service board and the second service board.
[0013] In another specific implementation, the first interface board belongs to an interface board set corresponding to the second service board, and the second service data includes service data transmitted between the third service board and the first service board.
[0014] In another specific implementation, the first interface board includes multiple interface boards, the first interface board belongs to an interface board set corresponding to the first service board, and the first interface board also belongs to an interface board set corresponding to the second service board. The second service data includes service data transmitted between the third service board and the fourth service board. The fourth service board includes a service board that does not correspond to the interface board set described by the first interface board. The interface board set corresponding to the fourth service board does not include the first interface board. The interface board set used by the fourth service board to transmit the second service data does not include the first interface board.
[0015] It should be understood that the first target central switch board set does not include a set of interface boards communicating with the first interface board. The first target central switch board set does not include the first central switch board.
[0016] It should be understood that in response to the failure of the first interface board, the first central switch board is not used to transmit the first service data, the first central switch board is also used to transmit the second service data, and the first central switch board is also used to transmit other service data. The service data bandwidth on the first central switch is not zero.
[0017] In combination with the first aspect, in certain implementations of the first aspect, the cluster system further includes N cross-board sets. The N cross-board sets correspond one-to-one to the N service boards, and each cross-board set includes at least one cross-board. Each service board transmits service data through the corresponding cross-board set and the corresponding interface board set. The first cross-board in the cross-board set corresponding to the first service board is used to transmit the first service data. The first cross-board is configured with at least one logical port. In response to a failure of the first interface board, the first cross-board is controlled to switch to the first logical port of the at least one logical port to execute the first policy.
[0018] Based on the above scheme, in response to the failure of any interface board, the present application controls the cross-board to switch different logical ports to execute different transmission strategies, so that the business data associated with the faulty board does not go to the central switching board connected to the faulty interface board. However, the business data transmitted between the central switching board and other interface boards in normal state is not affected, there is no need to isolate the entire switching plane, and the logical interface of the corresponding cross-board does not need to be switched. By configuring the transmission strategy of the virtual logical port in advance, fine-grained isolation of faults can be achieved by switching the logical ports of the cross-board, thereby improving the utilization rate of redundant resources and the operating efficiency of the overall cluster system.
[0019] It should be understood that the first cross-connect board includes one or more cross-connect boards, including cross-connect boards in the cross-connect board set corresponding to the first service board that participate in the transmission of the first service. Due to load balancing, the cross-connect boards in the cross-connect board set will evenly distribute the transmitted service data. When transmitting the first service data, the first cross-connect board will evenly distribute the first service data. The equalization steps for the first service by the first cross-connect board are determined by the transmission strategy. The equalization steps for the first service by the first cross-connect board are determined by the first target central switching board set.
[0020] In conjunction with the first aspect, in certain implementations of the first aspect, in response to a first interface board failure, it is detected that the total bandwidth of service data to be transmitted through the first target central switching board set exceeds a first threshold. Service data in the first service data that exceeds the first threshold is suspended. The first threshold includes an upper limit on the total bandwidth of the first target central switching board set.
[0021] Based on the above scheme, the present application responds to the failure of any interface board and controls the business data associated with the faulty board not to go to the central switching board connected to the faulty interface board. However, the business data transmitted between the central switching board and other interface boards in normal state is not affected, and there is no need to isolate the entire switching plane. By fine-tuning the isolation of faults, the utilization rate of redundant resources is improved, and the operating efficiency of the overall cluster system is also improved. During the switching process of business transmission, the bandwidth of business data that needs to be transmitted on the first target central switching board set is also monitored. When the business data after the fault switching exceeds the bandwidth upper limit, the transmission of the business data exceeding the upper limit is suspended. The solution provided by the present application also has the ability to control business issuance and recovery, and flexibly adjust the number of redundant resources according to the actual bandwidth of the business; when the number of faulty boards exceeds the protection capacity, the business bandwidth exceeding the protection capacity is actively recovered, thereby increasing the number of redundant resources, avoiding interruption of full-frame business, and reducing the impact of the fault.
[0022] In conjunction with the first aspect, some implementations of the first aspect further include receiving first information indicating a fault on the first interface board. Based on the first information, second information is sent instructing the first service board to transmit first service data based on a first policy.
[0023] Based on the above solution, the present application determines an interface board failure by receiving first information, and then uses second information to prevent service data associated with the failed board from being sent to the central switching board connected to the failed interface board. However, service data transmitted between the central switching board and other normal interface boards is not affected, eliminating the need to isolate the entire switching plane. This refined fault isolation improves the utilization of redundant resources and enhances the overall operational efficiency of the cluster system.
[0024] In combination with the first aspect, in some implementations of the first aspect, the second information is further used to instruct the first cross-connect board to switch to the first logical port.
[0025] Based on the above scheme, in response to a failure of any interface board, the present application controls the first cross-board to switch to the first logical port through the second information to execute the first transmission strategy, so that the business data associated with the faulty board does not go to the central switching board connected to the faulty interface board. However, the business data transmitted between the central switching board and other interface boards in normal state is not affected, there is no need to isolate the entire switching plane, and the logical interface of the corresponding cross-board does not need to be switched. By configuring the transmission strategy of the virtual logical port in advance, refined fault isolation is achieved by switching the logical port of the cross-board, the utilization rate of redundant resources is improved, and the operating efficiency of the overall cluster system is also improved.
[0026] In a second aspect, a cluster system is provided. The cluster system includes a control unit, N service boards, N interface board sets, and M central switching boards, where N is greater than or equal to 3 and M is greater than or equal to 2. Each interface board set includes at least one interface board, and there is a one-to-one correspondence between the N service boards and the N interface board sets. Each service board sends or receives service data through an interface board in its corresponding interface board set. Each central switching board is connected to at least three interface boards, and the at least three interface boards belong to different interface board sets. In response to a failure of a first interface board connected to a first central switching board in the M central switching boards, the control unit controls the first service board to transmit first service data to a second service board based on a first policy, and controls the third service board to transmit second service data based on a second policy. The first interface board belongs to the interface board set corresponding to the first service board or the second service board. The first policy is used to indicate a first target central switching board set for the first service data, and the second policy is used to indicate a second target central switching board set for the second service data. The first target central switching board set includes central switching boards connected to at least one interface board other than the first interface board in the interface board set to which the first interface board belongs, and the second target central switching board set includes the first central switching board.
[0027] Based on the above solution, in this application, the control unit responds to the failure of any interface board by preventing service data associated with the failed board from being sent to the central switching board connected to the failed interface board. However, service data transmitted between the central switching board and other normal interface boards remains unaffected, eliminating the need to isolate the entire switching plane. This refined fault isolation improves the utilization of redundant resources and enhances the overall operational efficiency of the cluster system.
[0028] In combination with the second aspect, in certain implementations of the second aspect, the cluster system further includes N cross-board sets. The N cross-board sets correspond one-to-one to the N service boards, each cross-board set includes at least one cross-board, and each service board transmits service data through the corresponding cross-board set and the corresponding interface board set. The first cross-board in the cross-board set corresponding to the first service board is used to transmit the first service data. The first cross-board is configured with at least one logical port. In response to a failure of the first interface board, the control unit controls the first cross-board to switch to the first logical port of the at least one logical port to execute the first policy.
[0029] Based on the above scheme, in response to a failure of any interface board, the control unit of the present application controls the cross-board to switch different logical ports to execute different transmission strategies, so that the business data associated with the faulty board does not go to the central switching board connected to the faulty interface board. However, the business data transmitted between the central switching board and other interface boards in normal state is not affected, there is no need to isolate the entire switching plane, and the logical interface of the corresponding cross-board does not need to be switched. By configuring the transmission strategy of the virtual logical port in advance, refined fault isolation is achieved by switching the logical port of the cross-board, the utilization rate of redundant resources is improved, and the operating efficiency of the overall cluster system is also improved.
[0030] In conjunction with the second aspect, in certain implementations of the second aspect, the cluster system further includes a bandwidth detection module. In response to a failure of the first interface board, the bandwidth detection module detects that the total bandwidth of service data to be transmitted through the first target central switching board set exceeds a first threshold. The control unit suspends service data in the first service data that exceeds the first threshold. The first threshold includes an upper limit on the total bandwidth of the first target central switching board set.
[0031] Based on the above solution, in the present application, in response to a failure of any interface board, the control unit controls the service data associated with the faulty board not to be sent to the central switching board connected to the faulty interface board. However, the service data transmitted between the central switching board and other interface boards in normal state is not affected, and there is no need to isolate the entire switching plane. By fine-tuning the fault isolation, the utilization rate of redundant resources is improved, and the operating efficiency of the overall cluster system is also improved. The cluster system of the present application also includes a bandwidth detection module. During the switching process of service transmission, the bandwidth detection module also monitors the bandwidth of service data that needs to be transmitted on the first target central switching board set. When the service data after the fault switching exceeds the bandwidth upper limit, the bandwidth detection module reports an alarm message to the control unit. The control unit suspends the transmission of the service data that exceeds the bandwidth upper limit based on the alarm message. The solution provided by the present application also has the ability to control service issuance and recovery, and flexibly adjusts the number of redundant resources according to the actual bandwidth of the service; when the number of faulty boards exceeds the protection capacity, the service bandwidth exceeding the protection capacity is actively recovered, thereby increasing the number of redundant resources, avoiding interruption of the entire frame service, and reducing the impact of the failure.
[0032] It should be understood that the information transmission between the bandwidth detection module and the control unit can be transmitted through separate information, or the alarm information can be carried in other information for reporting, and this application does not make any special restrictions on this.
[0033] It should also be understood that the information transmission between other modules or control units in the present application is similar to the information transmission mechanism between the bandwidth detection module and the control unit, and will not be repeated in this application.
[0034] In conjunction with the second aspect, in certain implementations of the second aspect, a control unit receives first information. The first information indicates a fault in a first interface board. The control unit sends second information based on the first information. The second information instructs the first service board to transmit first service data based on a first policy.
[0035] Based on the above solution, the control unit of the present application determines that an interface board is faulty by receiving first information and then sends second information indicating that service data associated with the faulty board is not transmitted to the central switching board connected to the faulty interface board. However, service data transmitted between the central switching board and other normal interface boards is not affected, eliminating the need to isolate the entire switching plane. This refined fault isolation improves the utilization of redundant resources and enhances the overall operational efficiency of the cluster system.
[0036] In combination with the second aspect, in some implementations of the second aspect, the second information is used to instruct the first cross-connect board to switch to the first logical port.
[0037] Based on the above solution, in the present application, in response to a failure of any interface board, the control unit sends a second message to instruct the first cross-board to switch to the first logical port to execute the first transmission strategy, so that the business data associated with the faulty board does not go to the central switching board connected to the faulty interface board. However, the business data transmitted between the central switching board and other interface boards in normal state is not affected, there is no need to isolate the entire switching plane, and the logical interface of the corresponding cross-board does not need to be switched. By configuring the transmission strategy of the virtual logical port in advance, refined fault isolation can be achieved by switching the logical port of the cross-board, thereby improving the utilization rate of redundant resources and the operating efficiency of the overall cluster system.
[0038] In conjunction with the second aspect, in certain implementations of the second aspect, the cluster system further includes a fault detection module that inspects each interface board in the cluster system according to a first time period and sends first information indicating a fault in the first interface board.
[0039] Based on the above scheme, the cluster system in the present application also includes a fault detection module. The fault detection module is used to inspect and detect whether there is a fault in the interface board. When the fault detection module detects a fault in the first interface board, it reports the first information. The control unit receives the first information to determine the fault in the interface board, and sends the second information to instruct the first cross-board to switch to the first logical port to execute the first transmission strategy, so that the business data associated with the faulty board does not go to the central switching board connected to the faulty interface board. However, the business data transmitted between the central switching board and other interface boards in normal state is not affected, there is no need to isolate the entire switching plane, and the logical interface of the corresponding cross-board does not need to be switched. By configuring the transmission strategy of the virtual logical port in advance, refined fault isolation is achieved by switching the logical port of the cross-board, the utilization rate of redundant resources is improved, and the operating efficiency of the overall cluster system is also improved.
[0040] In conjunction with the second aspect, in certain implementations of the second aspect, the first interface board includes a communication module. The first interface board reports first information via the communication module. The first information is used to indicate a fault in the first interface board.
[0041] Based on the above solution, the first interface board in the present application also includes a communication module. When the first interface board fails, it will automatically report the fault information through the communication module. When the first interface board fails, the first interface board reports the first information to the control unit. The control unit receives the first information to determine the interface board failure, and sends the second information to instruct the first cross-board to switch to the first logical port to execute the first transmission strategy, so that the business data associated with the faulty board does not go to the central switching board connected to the faulty interface board. However, the business data transmitted between the central switching board and other interface boards in normal state is not affected, there is no need to isolate the entire switching plane, and the logical interface of the corresponding cross-board does not need to be switched. By configuring the transmission strategy of the virtual logical port in advance, refined fault isolation is achieved by switching the logical port of the cross-board, the utilization rate of redundant resources is improved, and the operating efficiency of the overall cluster system is also improved.
[0042] In a third aspect, a data transmission device is provided. The device includes at least one processor configured to execute a computer program or instruction stored in at least one memory, so that the data transmission device performs the method of the first aspect or any possible implementation thereof.
[0043] In a fourth aspect, a computer-readable storage medium is provided, wherein computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed on a computer, the method of the first aspect or any possible implementation thereof is executed.
[0044] In a fifth aspect, a computer program product is provided. The computer program product includes computer executable code or computer executable instructions. When the computer executable code or computer executable instructions are executed by a computer, the method of the first aspect or any possible implementation thereof is performed. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A schematic diagram of a possible application scenario of an embodiment of the present application is shown.
[0046] Figure 2 A schematic diagram of a possible network device hardware structure is shown.
[0047] Figure 3 A possible OTN multi-chassis cluster system is shown.
[0048] Figure 4 One possible method of transmitting data is shown.
[0049] Figure 5 Another possible OTN multi-chassis cluster system is shown.
[0050] Figure 6 A schematic diagram showing a cross-connect board processing service flow is shown.
[0051] Figure 7 A possible data flow transmission diagram when a failure occurs in an OTN multi-chassis cluster system is shown.
[0052] Figure 8 A schematic diagram of the structure of an OTN device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0053] The technical solution in this application will be described below with reference to the accompanying drawings.
[0054] In order to facilitate understanding of the embodiments of the present application, the following explanations are provided.
[0055] First, in the text description or drawings of the embodiments of the present application shown below, the terms "first", "second", "third", "fourth", etc. and various numerical numbers are only used for the convenience of description and are not necessarily used to describe a specific order or sequence, and are not used to limit the scope of the embodiments of the present application.
[0056] Second, the terms "including" and "having" and any variations thereof in the embodiments of the present application shown below are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or apparatuses.
[0057] Third, in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. An embodiment or design described as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. The use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner to facilitate understanding.
[0058] Fourth, in the embodiments of the present application, a service refers to a service that can be carried by an optical transport network. For example, it can be an Ethernet service, a packet service, a wireless backhaul service, etc. Service data can also be referred to as a service signal, customer data, or customer service data.
[0059] The embodiments of the present application are applicable to optical networks, such as OTN. An OTN is generally composed of multiple optical switch network (OSN) devices connected by optical fibers, and can be formed into different topologies such as linear, ring, and mesh according to specific needs.
[0060] Figure 1 A schematic diagram of a possible application scenario of an embodiment of the present application is shown.
[0061] like Figure 1 The OTN 100 shown is composed of eight OSN devices 101, namely devices AH. 102 indicates an optical fiber used to connect two devices; 103 indicates a customer service interface used to receive or send customer service data. Figure 1 As shown, OTN 100 is used to transmit services for customer devices 1-3. Customer devices are connected to OSN devices through customer service interfaces. For example, Figure 1 In the example, customer devices 1-3 are connected to OSN devices A, H, and F respectively.
[0062] Depending on actual needs, an OSN device may have different functions. Generally speaking, OSN equipment is divided into optical layer equipment, electrical layer equipment, and optoelectronic hybrid equipment. Optical layer equipment refers to equipment that can process optical layer signals, such as optical amplifiers (OAs) and optical add-drop multiplexers (OADMs). OAs, also known as optical line amplifiers (OLAs), are primarily used to amplify optical signals to support longer transmission distances while ensuring the specific performance of the optical signal. OADMs are used to spatially transform optical signals so that they can be output from different output ports (sometimes also called directions). Electrical layer equipment refers to equipment that can process electrical layer signals, such as equipment that can process OTN signals. Optoelectronic hybrid equipment refers to equipment that has the ability to process both optical and electrical layer signals. It should be noted that, depending on specific integration needs, an OSN device can integrate multiple different functions.
[0063] Figure 2 A possible diagram of the hardware structure of a network device is shown below. For example, Figure 1 Device A in the figure. Specifically, OSN equipment 200 includes a service board 201, a cross-connect board 202, an interface board 203, an optical layer processing board (not shown), and a system control and communication board 204. Depending on specific needs, the type and number of boards included in a network device may vary. For example, a network device serving as a core node may not have a service board 201. Another example is a network device serving as an edge node that may have multiple service boards 201 or no optical cross-connect board 202. Another example is a network device that only supports electrical layer functions may not have an optical layer processing board.
[0064] The service board 201, cross-connect board 202, and interface board 203 are used to process OTN electrical layer signals. The service board 201 is used to receive and transmit various customer services, such as SDH services, packet services, Ethernet services, and fronthaul services. Furthermore, the service board 201 can be divided into client-side and line-side signal processors. The client-side optical transceiver module, also known as an optical transceiver, is used to receive and / or transmit service data. The signal processor is used to map and demap service data into data frames. The cross-connect board 202 is used to exchange data frames, completing the exchange of one or more types of data frames. The interface board 203 primarily processes the out-of-frame service interface of the cluster system. The signal processor is used to multiplex and demultiplex, or map and demap, data frames on the client and line sides. The system control and communication board 204 is used to implement system control and communication. Specifically, it can collect information from different boards or send control instructions to the corresponding boards. It should be noted that, unless otherwise specified, the specific components (such as signal processors) can be one or more, and this application does not impose any restrictions. It should also be noted that this application does not impose any restrictions on the type of single board included in the device, as well as the functional design and quantity of the single board. It should be noted that in a specific implementation, the above two single boards may also be designed as one single board. In addition, the network equipment may also include a power supply for backup, a fan for heat dissipation, etc.
[0065] With the development of informatization and cloud computing, demand for dedicated lines and video services is growing rapidly. These services are characterized by low bandwidth and high volume, requiring simple, fast, and flexible bandwidth adjustments. Optical transport networks (OTNs) are widely deployed in backbone lines, metropolitan cores, and metropolitan edge networks, offering inherent advantages of high quality, large capacity, and wide coverage.
[0066] Currently, the existing multi-chassis cluster system adopts a fixed backup solution and cannot dynamically adjust redundant resources according to the actual bandwidth of the business. Figure 3 The multi-chassis cluster system shown.
[0067] Figure 3 This diagram illustrates a possible OTN multi-chassis cluster system. The service interactions performed by the service boards of each OSN device can be abstractly divided into a service board. It should be understood that the OTN multi-chassis cluster system may include multiple OSN devices, each of which may include one or more service boards. This application does not impose any specific limitations on this.
[0068] This application divides the physical architecture of each business board and its cross-board and interface board into a business frame. The number of cross-boards and interface boards in each business frame is not specifically limited in this application. The number of cross-boards and interface boards between different business frames can be the same or different, and this application does not specifically limit this. Figure 3 The cluster architecture shown is used as a specific implementation method to describe the business interactions between the business boxes of this application.
[0069] In a service frame, service boards are connected to all cross-connect boards via a high-speed backplane bus. All cross-connect boards are fully connected to all interface boards via the backplane high-speed bus. Interface boards in different service frames are connected via a central switching board. At least one interface board in each service frame is connected to at least one central switching board, forming a switching plane. Each interface board in a service frame is connected to the central switching board in its switching plane, and there is no interconnection between switching planes. This application does not impose any specific restrictions on the number of service boards or service frames.
[0070] It should be understood that fully connecting all cross-connect boards and all interface boards means that every cross-connect board and every interface board is connected via a high-speed backplane bus. No interconnection between switching planes means that each central switching board is connected only to one interface board in each service frame, each interface board is connected only to the central switching board in its switching plane, each interface board belongs to only one switching plane, and each central switching board belongs to only one switching plane.
[0071] exist Figure 3 In a specific implementation shown, two service blocks are included: service block 1 and service block 2. Each service block includes a service board, four cross-connect boards, and four interface boards. Within each service block, the service board is connected to all cross-connect boards via a high-speed backplane bus, and all cross-connect boards are fully connected to all interface boards via the backplane high-speed bus. In addition to the two service blocks, four central switching boards are also included. The central switching board is connected to an interface board in each service block, forming a switching plane. For example, central switching board 1 is connected to interface board 1 in service block 1 and also to interface board 5 in service block 2. Interface board 1, central switching board 1, and interface board 5 form switching plane 1. For another example, central switching board 4 is connected to interface board 4 in service block 1 and also to interface board 8 in service block 2. Interface board 4, central switching board 4, and interface board 8 form switching plane 4. Similarly, the connections between switching planes 2 and 3 are not further described here. The service blocks employ a parallel interconnect architecture, with services forwarded between them based on the switching plane.
[0072] It should be understood that parallel interconnection architectures between chassis include both horizontal and vertical connection architectures. In a horizontal connection architecture, the interface board in service chassis 1 connects to the central switching board, which in turn connects to the interface board in service chassis 2, forming a switching plane. In a vertical connection architecture, there is no interconnection between the switching planes, and no optical fiber connection between the switching planes; instead, the switching planes are parallel.
[0073] In a specific implementation, the number of interface boards in a business frame is the same as the number of central switching boards, each interface board corresponds to each central switching board one-to-one, each interface board is connected to a different central switching board, and each central switching board is connected to a different interface board.
[0074] exist Figure 3 In the multi-chassis cluster system shown, if an interface board in a service frame fails, the switching plane where the failed interface board resides is isolated. This means that services transmitted by interface boards in other service frames within that switching plane are also isolated. As a non-limiting example, if interface board 5 in service frame 2 fails, the entire switching plane 1 is isolated, and the transmission links connecting interface board 1 in service frame 1 to switching plane 1 are also isolated. In other words, none of the four cross-connect boards in service frame 1 send service flows to switching plane 1, and service flows are transmitted only through the remaining three switching planes.
[0075] If there are still more service frames, and any interface board or central switch board in a switching plane fails, the entire switching plane is isolated. That is, the transmission links between the central switch board and all interface boards in the switching plane are isolated, and the cross-connect boards in the service frame no longer send service flows to the failed switching plane. All service flows are transmitted only through the remaining switching planes.
[0076] Based on the above technical solution, a multi-chassis cluster uses a switching plane to forward services. If an interface board or central switching board fails, the system isolates the entire switching plane, interrupting all services on that switching plane and requiring them to be transferred to other switching planes. Because multi-chassis clusters define bandwidth specifications based on devices, even if a single interface board or central switching board fails, service flows can still be transmitted through the remaining switching planes. However, if two points of failure occur—that is, an interface board or central switching board in another switching plane also fails—there are no redundant resources to share service flows, resulting in service interruption for the entire chassis.
[0077] This solution cannot dynamically adjust redundant resources based on the actual bandwidth of the business. When two or more points of failure occur, it cannot provide protection capabilities and the business of the entire frame will be interrupted. The granularity of redundant resource allocation is large and the system availability is low, which cannot provide sufficient availability and protection capabilities.
[0078] However, the multi-chassis cluster architecture creates a conflict between increasing the capacity and redundancy of the overall communication system. System capacity is constrained by process technology and integration levels. Given the same device capacity, it's impossible to simultaneously increase both capacity and redundancy, and the availability of fault redundancy protection solutions is low.
[0079] Therefore, it is necessary to reasonably allocate the system's redundant resources without increasing the size of the equipment or reducing the system capacity, reduce the risk of business interruption, and improve the system's availability.
[0080] This application proposes a data transmission method with refined control over fault isolation. When a fault point occurs in the system, the switching plane is dynamically allocated and adjusted based on the service traffic load. Refined isolation operations can be performed on the backup switching plane, that is, redundant resources are finely adjusted according to the severity of the fault, thereby maximizing the utilization of redundant resources and improving system availability and stability. Furthermore, this application proposes a method for controlling service degradation. When system resources are tight, this method prioritizes the stable operation of important services and automatically reduces the performance of less critical services, thereby ensuring the availability and stability of the entire system.
[0081] Below through Figure 4 It should be understood that the embodiment shown in the drawings is only a specific implementation method and does not constitute any limitation on the scope of protection of this application.
[0082] Figure 4 A possible data transmission method is shown, which is executed in a cluster system.
[0083] Figure 4 The cluster system shown includes at least three service boards, at least three interface board sets, and at least two central switching boards. Each interface board set includes at least one interface board. The three service boards correspond one-to-one to the three interface board sets. Each service board sends or receives service data through an interface board in its corresponding interface board set.
[0084] Each central switching board in the cluster system is connected to at least three interface boards, and the at least three interface boards belong to different interface board sets.
[0085] In one specific implementation, central switching board 1 is connected to interface board 1, interface board 2, and interface board 3, respectively. Interface board 1, interface board 2, and interface board 3 each belong to a different interface board set. At least one interface board in each interface board set is connected to central switching board 1, forming a switching plane. At least one interface board in each interface board set, other than the one connected to central switching board 1, is connected to central switching board 2, forming another switching plane. No service data is exchanged between the different switching planes, and the different switching planes are parallel.
[0086] It should be understood that a central switching board can also be connected to two or more interface boards in the interface board set to form a switching plane; an interface board can also be connected to multiple central switching boards to form a switching plane; and there is no exchange of service data between different switching planes. This application does not impose any specific restrictions on the number of central switching boards and connected interface boards.
[0087] When all interface boards in the cluster system are normal, service board 1 transmits service data 1 between service board 2 through interface board assembly 1 and the central switching board to which the interface board is connected.
[0088] It should be understood that business data 1 is a specific implementation method of the first business data in the implementation method of this application, and this application does not make any special limitations on this.
[0089] It should be understood that business board 1 is a specific implementation of the first business board in the implementation of this application, business board 2 is a specific implementation of the second business board in the implementation of this application, and business board 3 is a specific implementation of the third business board in the implementation of this application. This application does not make any special limitations on this.
[0090] It should be understood that the central switching board 1 is a specific implementation of the first central switching board in the implementation of this application, and this application does not make any special limitation on this.
[0091] It should be understood that interface board set 1 is a specific implementation of the first interface board set in the implementation of this application, and interface board set 2 is a specific implementation of the second interface board set in the implementation of this application. This application does not make any special restrictions on this.
[0092] The cluster system also includes at least 3 cross-board sets, which correspond to 3 service boards one by one. Each cross-board set includes at least one cross-board, and each service board transmits service data through the corresponding cross-board set and the corresponding interface board set.
[0093] When all interface boards in the cluster system are functioning properly, service board 1 transmits service data 1 between service board 2 and cross-connect board assembly 1, interface board assembly 1, and the central switching board to which the interface boards are connected. It should be understood that service board 2 receives service data 1 through interface board assembly 2 and cross-connect board assembly 1.
[0094] When all interface boards in a cluster system are working properly, to ensure load balancing, service data sent by a service board is distributed evenly across all cross-connect boards in the cross-connect board set that carry that service. The cross-connect boards then transmit the carried service data to all connected interface boards for second-level distribution. Each interface board then transmits the carried service data to another service board via the connected switching plane.
[0095] As a specific implementation, all cross-connect boards in cross-connect board set 1 and all interface boards in interface board set 1 are operating normally and participate in the transmission of service data 1. Service board 1 sends service data 1, and to ensure load balancing, service data 1 is evenly distributed across all cross-connect boards in cross-connect board set 1. The service data carried by cross-connect board 1 is transmitted to interface board set 1 and then evenly distributed twice by interface board 1 and interface board 4. The service data carried by cross-connect board 4 is also transmitted to interface board set 1 and then evenly distributed twice by interface board 1 and interface board 4. Interface board 1 transmits the service data, which has been evenly distributed twice by cross-connect board 1 and cross-connect board 4, to the connected central switching board 1. Central switching board 1 transmits the service data to interface board 2. Similarly, interface board 4 transmits the service data, which has been evenly distributed twice by cross-connect board 1 and cross-connect board 4, to the connected central switching board 2. Central switching board 2 transmits the service data to interface board 5. Interface board set 2 then transmits the received data to service board 2 via cross-connect board set 2.
[0096] It should be understood that when data is transmitted between service boards, the service data will be evenly distributed for load balancing, and this application does not specifically limit the load balancing method.
[0097] It should be understood that the transmission method and steps of business data 2 transmitted between business board 3 and business board 1 are the same as those of business data 1, and the transmission method and steps of business data 3 transmitted between business board 3 and business board 2 are the same as those of business data 1. This application will not go into details here.
[0098] In response to a failure of a first interface board connected to central switching board 1, service board 1 is controlled to transmit first service data to service board 2 based on a first policy, and service board 3 is controlled to transmit second service data based on a second policy. The first interface board belongs to an interface board set corresponding to service board 1 or service board 2. The first policy is used to indicate a first target central switching board set to transmit service data 1, and the second policy is used to indicate a second target central switching board set to transmit the second service data. The first target central switching board set includes central switching boards that are connected to at least one remaining interface board in the interface board set to which the first interface board belongs, excluding the first interface board. The second target central switching board set includes central switching board 1.
[0099] Cross-connect board 1 in cross-connect board set 1 corresponding to service board 1 is used to transmit service data 1. Cross-connect board 1 is configured with at least one logical port. In response to a failure of the first interface board, cross-connect board 1 switches to the first logical port to execute the first policy.
[0100] It should be understood that the cross-board set 1 is a specific implementation of the first cross-board set in the implementation of this application, and the cross-board 1 is a specific implementation of the first cross-board in the implementation of this application, and this application does not make any special restrictions on this.
[0101] It should be understood that the cross-connect board is configured with virtual logical ports, and each logical port can be configured with different policies. Different policies are used to indicate which central switching boards and switching planes to transmit service data through.
[0102] In a specific implementation, interface board 1 is a specific implementation of a first interface board. Interface board 1 is connected to central switching board 1, and interface board 1 belongs to interface board set 1. In response to a failure of interface board 1, service board 1 transmits service data 1 to service board 2 based on a first strategy, and service board 3 transmits service data 3 based on a second strategy. The first target central switching board set includes central switching boards connected to at least one interface board other than interface board 1 in interface board set 1. The first target central switching board set includes central switching board 2 connected to interface board 4. The second target central switching board set includes central switching board 1. Due to the requirement of load balancing during transmission, the second target central switching board set includes central switching board 1 and central switching board 2.
[0103] It should be understood that in the above specific implementation, business data 1 is a specific implementation of the first business data, and business data 3 is a specific implementation of the second business data. This application does not make any special restrictions on this.
[0104] In another specific implementation, interface board 2 is a specific implementation of the first interface board. Interface board 2 is connected to central switching board 1 and belongs to interface board set 2. In response to a failure of interface board 2, service board 1 transmits service data 1 to service board 2 based on a first policy, and service board 3 transmits service data 2 based on a second policy. The first target central switching board set includes central switching boards connected to at least one interface board in interface board set 2 other than interface board 2. The first target central switching board set includes central switching board 2 connected to interface board 5. The second target central switching board set includes central switching board 1. Due to the requirement of load balancing during transmission, the second target central switching board set includes central switching board 1 and central switching board 2.
[0105] It should be understood that in the above specific implementation, business data 1 is a specific implementation of the first business data, and business data 2 is a specific implementation of the second business data. This application does not make any special restrictions on this.
[0106] It should be understood that the first interface board includes one or more interface boards, and this application does not impose any special limitation on this.
[0107] In another specific implementation, a first interface board includes interface board 1 and interface board 2. The combination of interface board 1 and interface board 2 serves as a specific implementation of the first interface board. Interface board 1 and interface board 2 are both connected to central switching board 1. Interface board 1 belongs to interface board set 1, and interface board 2 belongs to interface board set 2. In response to a failure of both interface board 1 and interface board 2, service board 1 transmits service data 1 to service board 2 based on a first policy, and service board 3 transmits service data 2 based on a second policy. The first target central switching board set includes central switching boards connected to at least one interface board in interface board set 1 other than interface board 1, and / or includes central switching boards connected to at least one interface board in interface board set 2 other than interface board 2. The first target central switching board set includes central switching board 2 connected to interface board 4, and / or central switching board 2 connected to interface board 5. The second target central switching board set includes central switching board 1. Due to the requirement for load balancing during transmission, the second target central switching board set includes central switching board 1 and central switching board 2.
[0108] It should be understood that in the above-mentioned specific implementation method, business data 1 is a specific implementation method of the first business data, and the second business data includes the business transmitted between business board 3 and other business boards other than business board 1 and business board 2 (not drawn in the figure). This application does not make any special limitations on this.
[0109] It should be understood that the number of service boards, the number of central switching boards, and the number of boards in each set in the accompanying drawings are only schematically drawn, which is only a specific implementation method and does not constitute any limitation on the scope of protection of this application.
[0110] This technology solves service data transmission issues by switching service data from faulty lines to healthy lines. Compared to traditional technologies, the isolation granularity of faulty lines is reduced from the entire switching plane to a single interface board link. The functioning central switching board can still transmit other normal services. This more rational and efficient use of redundant resources improves the granularity of redundancy adjustment and enhances overall system efficiency.
[0111] In response to a first interface board failure, the total bandwidth of service data to be transmitted through the first target central switching board set is detected. When the total bandwidth of the transmitted service data exceeds a first threshold, the transmission of the portion of the first service data that exceeds the first threshold is suspended. The first threshold includes an upper limit on the total bandwidth of the first target central switching board set.
[0112] It should be understood that when switching transmission strategies, the corresponding central switching board in the target central switching board set will change. The central switching board will continue to transmit the failed-over service while ensuring the normal transmission of the original service. If the service data to be transmitted exceeds the total bandwidth limit after the failover, the service data exceeding the bandwidth limit will be suspended to ensure smooth transmission of the entire service.
[0113] The cluster system determines that a first interface board has failed by receiving first information. The first information indicates the failure of the first interface board. A control unit of the cluster system sends second information based on the first information. The second information instructs the first service board to transmit service data 1 based on a first policy. The second information instructs cross-connect board 1 to switch to a first logical port.
[0114] In a specific implementation manner, the first information is reported automatically by the fault interface board.
[0115] In another specific implementation, the cluster system further includes a fault detection module that performs inspections on each module and board in the system according to a specific time period, and reports a fault via the first information when a fault is found.
[0116] Figure 5 This illustrates another possible OTN multi-chassis cluster system. The service interactions performed by the service boards of each OSN device can be abstractly divided into a service frame. It should be understood that an OTN system can include multiple OSN devices, each of which can have one or more service boards. This application does not impose any specific limitations on this.
[0117] It should be understood that the business frame includes business boards, and the services transmitted by the business boards include uplink services, downlink services, input services, output services, etc. In some implementations, it is also called an input / output frame (input / output frame, IO frame). This application does not make any special restrictions on this.
[0118] This application divides the physical architecture of each business board and its cross-board and interface board into a business frame. The number of cross-boards and interface boards in each business frame is not specifically limited in this application. The number of cross-boards and interface boards between different business frames can be the same or different, and this application does not specifically limit this. Figure 5 The multi-frame cluster architecture shown is used as a specific implementation method to describe the service interaction between the service frames of this application.
[0119] In each service frame, service boards are connected to all cross-connect boards via the backplane high-speed bus. All cross-connect boards are fully connected to all interface boards via the backplane high-speed bus. Interface boards in different service frames are connected via the central switching board. An interface board in each service frame is connected to a central switching board via the backplane high-speed bus, forming a switching plane. Each interface board in each service frame is connected to only one central switching board; there is no interconnection between switching planes.
[0120] Figure 5 The figure shows four service boxes, each containing a service board. Each service board is connected to four cross-connect boards via a high-speed backplane bus. The four cross-connect boards and four interface boards in a service box are fully connected via the high-speed backplane bus. Each interface board is connected to a central switching board, forming four switching planes. Service data is sent from a service board in a service box, transmitted to the central switching board via the cross-connect board and interface board, and then forwarded to another service box via the switching plane containing the central switching board. The interface board in the other service box receives the service data and then transmits it to the service board via the cross-connect board.
[0121] In a specific data transmission process, service board 1 in service block 1 sends service flow Flow1 to service board 2. The service data passes through the cross-connect board and interface board in service block 1 and is forwarded to service board 2 in service block 2 via the central switch board in switching plane 0 / 1 / 2 / 3. Service board 1 in service block 1 sends service flow Flow2 to service board 3. The service data passes through the cross-connect board and interface board in service block 1 and is forwarded to service board 3 in service block 3 via the central switch board in switching plane 0 / 1 / 2 / 3.
[0122] It should be understood that service flow Flow 1 includes the transmission path for a service transmitted from service block 1 to service block 2 via switching planes 0 / 1 / 2 / 3, and Flow 2 includes the transmission path for a service transmitted from service block 1 to service block 3 via switching planes 0 / 1 / 2 / 3. There is no limitation on the number of services transmitted through service flows Flow 1 and Flow 2, the bandwidth of the transmission paths, or the data size.
[0123] It should be understood that business flow Flow 1 is a specific implementation method of the first business data in the implementation method of this application and should not constitute any limitation on the scope of protection of this application.
[0124] It should be understood that transmitting services through switching planes 0 / 1 / 2 / 3 includes evenly processing service flows through switching planes 0, 1, 2, and 3, thereby achieving load balancing.
[0125] It should be understood that the switching plane refers to a combination of central switching boards and interface boards in an interconnected relationship. Each switching plane includes at least one central switching board and at least one interface board in each service frame.
[0126] Figure 6 A schematic diagram showing a cross-connect board processing service flow is shown.
[0127] During normal operation of the entire ONT multi-chassis cluster system, consider Service Board 1 in Service Board 1. Service Board 1 in Service Board 1 sends service flows Flow 1 and Flow 2 to the cross-connect board. After receiving the service flows, the cross-connect board determines the corresponding port numbers. In one specific implementation, the service flows and their corresponding port numbers are shown in Table 1.
[0128] Table 1
[0129] Business flow number Port number Flow 1 Port 0 Flow 2 Port 0 —— —— —— ——
[0130] As shown in Table 1, both Flow 1 and Flow 2 are configured on Port 0. The cross-connect board then determines the switching plane link to transmit the service flow based on the port number corresponding to the service flow. In one specific implementation, the correspondence between port numbers and switching plane links is shown in Table 2.
[0131] Table 2
[0132] Port number Switching plane link Port 0 Switching plane 0 / 1 / 2 / 3 links Port 1 Switching plane 1 / 2 / 3 links Port 2 Switching plane 0 / 2 / 3 link —— ——
[0133] As shown in Table 2, links to switching planes 0 / 1 / 2 / 3 are configured for Port 0. That is, when the multi-chassis cluster system is operating normally, service flows Flow 1 and Flow 2 are sent from service chassis 1 to switching planes 0 / 1 / 2 / 3. These flows are evenly distributed and processed by switching planes 0, 1, 2, and 3, achieving load balancing.
[0134] It should be understood that the number of service flows sent by a service frame, the ports configured for each service flow, and the switching plane links configured for each port are described in this embodiment as an optional implementation method and do not constitute any limitation on this application. When an actual cluster system runs cross-frame services, service flows, ports, and switching plane links can be configured based on the actual services, and this application does not impose any special limitations on this.
[0135] When the business frame is in normal operation, the fault detection module is used to inspect the working status of each device in the cross-system to inspect whether the cross-board, interface board, and central switching board in each business frame are in normal operation. If a fault is found in the interface board or central switching board in a business frame, it will be reported to the control unit. Since there is a faulty board in the system, there will be business flows that are forced to be interrupted due to the faulty board. After receiving the fault information, the control unit is used to send an indication message to the cross-board to which the forced interrupted business flows, instructing the cross-board to switch to the logical port. Among them, the switching plane used by the switched logical port does not include the switching plane where the faulty board is located. Only by changing the logical interface of the forced interrupted business flow, it is switched to other switching planes for business transmission; other business flows that are not affected by the faulty board continue to be transmitted through the originally configured ports and are not affected by the faulty board. The following is Figure 7 The position of the faulty board shown is used as a specific implementation method to describe the transmission of data streams during a fault, and does not constitute any limitation on the scope of protection of this application.
[0136] It should be understood that the fault detection module is a specific implementation of the fault detection module in the implementation of this application and does not constitute any limitation on the protection scope of this application.
[0137] Figure 7 A possible data flow transmission diagram when a failure occurs in an OTN multi-chassis cluster system is shown.
[0138] In a specific implementation, the fault information is reported to the control unit by the fault detection module when a faulty board is detected. During the inspection process, the fault detection module detects that the interface board connected to the switching plane 0 in the service box 2 is faulty (see FIG. Figure 7 The fault detection module reports the fault information of the interface board to the control unit.
[0139] In another specific implementation, the fault information can also be reported to the control unit by the faulty board. If a service failure occurs on the interface board in service box 2 connected to switching plane 0 and the service cannot be transmitted normally, the interface board will report the fault information to the control unit.
[0140] It should be understood that the fault board is a specific implementation of the first interface board in the implementation of this application and does not constitute any limitation on the protection scope of this application.
[0141] It should be understood that the fault information is a specific implementation of the first information in the implementation of this application and does not constitute any limitation on the protection scope of this application.
[0142] After receiving the fault information, the control unit determines the affected service flows. Upon receiving the fault information for the interface board, the control unit determines that the affected service flows include Flow 1, which is sent from service block 1 to service block 2. Because the interface board failure interrupts transmission, the control unit switches the service bandwidth from service block 1 to service block 2 to switching planes 1 / 2 / 3. Without exceeding the total bandwidth of switching planes 1 / 2 / 3, the control unit switches the port corresponding to Flow 1 in Table 1 from Port 0 to logical port 1, as shown in Table 3.
[0143] Table 3
[0144] Business flow number Port number Flow 1 Port 1 Flow 2 Port 0 —— —— —— ——
[0145] Based on the switching plane link information pre-configured for each port in Table 2, the control unit switches the transmission link of Flow 1 from switching plane 0 to switching planes 1 / 2 / 3, and evenly distributes the service flow among switching planes 1, 2, and 3.
[0146] It should be understood that the switching plane link information is a specific implementation of the service flow transmission strategy in the implementation of this application, and does not constitute any limitation on the protection scope of this application.
[0147] Service flow Flow 1 is not sent to the faulty switching plane 0. By switching to the logical port, service flow Flow 1 is restored to normal. In addition to service flow Flow 1 sent to service frame 2, which is transmitted through the remaining three switching planes (switching planes 1 / 2 / 3), service flows sent to other service boards are still transmitted through the four switching planes (switching planes 0 / 1 / 2 / 3). There is no need to isolate the transmission services on the entire switching plane 0. In other words, the data flow of inter-frame services sent to other service frames still corresponds to the transmission port Port 0, and no plane linkage switching is required.
[0148] When fault isolation occurs in the system, redundant resources are used. To ensure system availability, the bandwidth of new service flow Flow 1 must not exceed the total bandwidth of switching planes 1, 2, and 3. The system can also adjust the service bandwidth cap to limit new services that exceed the bandwidth cap, ensuring the normal transmission of existing service flows. After the fault is resolved, the redundant resources are released, and the service bandwidth cap returns to its original value, enabling dynamic adjustment of redundant resources.
[0149] Based on the technical solution of the present application, the hardware device provides the ability to distribute services based on ports, supports the configuration of multiple logical ports on the physical ports between frames, and each logical port can be flexibly configured to go to the switching plane. In a multi-frame cluster scenario, if an interface board failure occurs in any business frame, the cross-board control of each business frame prevents the cross-frame services associated with the faulty frame from going to the faulty interface board, and other unrelated cross-frame services are not affected. Fine-grained fault isolation without isolating the entire switching plane improves the utilization of redundant resources. At the same time, the system also has the ability to control service issuance and flexibly adjust the number of redundant resources according to the actual bandwidth of the service. The system also has the ability to control service recovery. When the number of faulty boards exceeds the protection capacity, it actively recovers the service bandwidth that exceeds the protection capacity, thereby increasing the number of redundant resources, avoiding interruption of services in the entire frame, and reducing the impact of the failure.
[0150] In another specific implementation, a multi-chassis cluster system experiences two faults. In addition to the fault of the interface board in service chassis 2 that connects to switching plane 0, the fault also occurs in the interface board in service chassis 3 that connects to switching plane 1. The fault information of these interface boards can be reported to the control unit by the faulty board itself, or by a fault detection module, which is not specifically limited in this application.
[0151] After receiving the fault information, the control unit determines the affected service flows. In addition to service flow Flow 1, service flow Flow 2, which is sent from service block 1 to service block 3, is also affected. The control unit's adjustments to service flow Flow 1 are not detailed here. The control unit also switches the service bandwidth from service block 1 to service block 3 to switching plane 0 / 2 / 3. Without exceeding the total bandwidth of switching plane 0 / 2 / 3, the port corresponding to Flow 2 in Table 1 is switched from Port 0 to logical port Port 2, as shown in Table 4.
[0152] Table 4
[0153] Business flow number Port number Flow 1 Port 1 Flow 2 Port 2 —— —— —— ——
[0154] Based on the switching plane link information pre-configured for each port in Table 2, the control unit switches the transmission link of Flow 2 from switching plane 1 to switching planes 0 / 2 / 3, and evenly distributes the service flow among switching planes 0, 2, and 3.
[0155] Service flow Flow 2 is not sent to the faulty switching plane 1. By switching to the logical port, service flow Flow 2 is restored to normal. Except for service flow Flow 2 sent to service frame 3 and service flow Flow 1 sent to service frame 2, inter-frame service flows sent to other service frames continue to be transmitted normally through switching planes 0 / 1 / 2 / 3. There is no need to isolate the transmission services on switching planes 0 and 1. In other words, the transmission port corresponding to inter-frame service flows sent to other service frames remains Port 0, and no plane linkage switching is required.
[0156] It should be understood that when performing service switching of service frame 3, the system needs to adjust the service bandwidth upper limit. The specific steps and principles are the same as the above implementation method, and this application will not repeat them here.
[0157] In one specific implementation, service flow Flow 1 is switched from switching plane 0 / 1 / 2 / 3 to the transmission link of switching plane 1 / 2 / 3, and service flow Flow 2 is switched from switching plane 0 / 1 / 2 / 3 to the transmission link of switching plane 0 / 2 / 3, achieving two-point failover. Compared with existing solutions, this solution does not isolate the entire switching plane. Instead, it adjusts the logical ports to transmit service flows through links in other switching planes. Surviving links in the switching plane continue to transmit service data normally, while only the link between the faulty interface board and the central switching board is isolated. This maximizes the use of redundant resources and enables dynamic adjustment of redundant resources.
[0158] The technical solution of this application solves the existing problem of the inability to dynamically adjust redundant resources in the backup switching plane. It also addresses the problem of global service interruptions caused by insufficient redundant resources. By dynamically allocating and adjusting redundant resources based on the actual bandwidth requirements of the service and the service traffic load, this solution can maximize system availability and stability, resolving existing issues such as the inability to simultaneously isolate multiple points of failure, low redundant resource utilization, and global service interruptions.
[0159] It should be understood that Figure 7 The faulty board shown is only one specific fault mode. The data transmission method used when other boards fail is the same, and this application will not go into details here.
[0160] It should be understood that the specific logical port configuration after the faulty board occurs and the switching plane link configured for the logical port in the above implementation are only an optional implementation, and this application does not impose any special restrictions on this.
[0161] It should be understood that when three-point failures or multiple-point failures occur in a multi-frame cluster system, the steps and principles for adjusting redundant resources are the same as the above-mentioned implementation methods, provided that the total bandwidth of the switching plane is not exceeded. This application will not repeat them here, and it should not be considered to exceed the scope of protection of this application.
[0162] In another specific implementation, the control unit receives service failure information from an interface board in service block 2 connected to switching plane 0, necessitating the switching plane switchover of service flow Flow 1 from switching plane 0 / 1 / 2 / 3 to switching plane 1 / 2 / 3. When the service bandwidth from service block 1 to service block 2 is switched to switching plane 1 / 2 / 3, it exceeds the total bandwidth of switching planes 1 / 2 / 3. The control unit then proactively reclaims the portion of the service from service block 1 to service block 2 that exceeds the total bandwidth of switching planes 1 / 2 / 3 from the device side, reducing the service bandwidth to below the total bandwidth of switching planes 1 / 2 / 3. The control unit then switches the port number corresponding to service flow Flow 1 from Port 0 to logical port Port 1, preventing service flow 1 from being sent to the faulty switching plane 0 and restoring normal transmission of service flow 1. Services between other frames, except service block 2, continue to be transmitted through Port 0, without the need for plane-linked switching.
[0163] When a fault is isolated in the system and redundant resources are occupied, to maintain system availability, new services must not exceed the total bandwidth of the remaining switching planes 1 / 2 / 3. The system adjusts the service bandwidth cap to limit the number of new services that exceed the bandwidth cap. After the fault is resolved, the redundant resources are released, the service bandwidth cap is restored to its original value, and the services that were reclaimed on the device side are reissued.
[0164] Based on the technical solution of the present application, the hardware device provides the ability to distribute services based on ports, supports the configuration of multiple logical ports on the physical ports between frames, and each logical port can be flexibly configured to go to the switching plane. In a multi-frame cluster scenario, if an interface board failure occurs in any service frame, the cross-board control of each service frame prevents the cross-frame services associated with the faulty frame from going to the faulty interface board, and other unrelated cross-frame services are not affected. Fine-grained fault isolation without isolating the entire switching plane improves the utilization rate of redundant resources. At the same time, the system also has the ability to control service issuance and flexibly adjust the number of redundant resources according to the actual bandwidth of the service. The system also has the ability to control service recovery. When the service bandwidth that needs to be switched exceeds the protection capacity, it actively recovers the service bandwidth that exceeds the protection capacity, thereby increasing the number of redundant resources, avoiding interruption of services in the entire frame, and reducing the impact of the failure.
[0165] It should be understood that when a two-point failure, a three-point failure, or a multi-point failure occurs, if the service bandwidth of the switching plane that needs to be switched exceeds the total bandwidth of the switching plane, the principles and steps for actively recovering the services that exceed the bandwidth are the same as the above-mentioned implementation method. This application will not repeat them here, and it should not be considered to exceed the scope of protection of this application.
[0166] Figure 8 FIG. 1 shows a schematic diagram of the structure of an OTN device provided in an embodiment of the present application. Figure 8 As shown, the OTN device 800 includes a processor 810 and a transceiver 820 .
[0167] During implementation, each step of the processing flow may be performed by hardware integrated logic circuits or software instructions in the processor 810 to complete the method described in the above implementation.
[0168] In some specific implementations, the control unit 801 includes a processor 810 and a transceiver 820 . The transceiver 820 is used to receive information and transmit it to the processor 810 . The transceiver 820 is also used to send information generated by the processor 810 .
[0169] In some other specific implementations, the control unit 801 includes a processor 810, a transceiver 820, and a memory 830. The processor 810 is configured to execute a program stored in the memory 830 to implement the above implementations.
[0170] In some specific implementations, the OTN device 800 further includes a fault detection module 840 , which is configured to inspect whether various components in the OTN device are operating normally, and report fault information when a fault is detected.
[0171] In some specific implementations, the OTN device 800 further includes a bandwidth detection module 850, which is used to detect the service bandwidth required to be transmitted by each transmission module in the system and report information when the bandwidth required to transmit the service exceeds the module bandwidth upper limit.
[0172] In the embodiments of the present application, the processor 810 can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application can be directly implemented as a hardware processor, or can be implemented by a combination of hardware and software units in the processor.
[0173] In addition, the OTN device 800 may include one or more processors 810 .
[0174] Optionally, the OTN device may further include a memory 830, wherein the program code executed by the processor 810 to implement the above method may be stored in the memory 830. The OTN device 800 may include one or more memories 830.
[0175] Specifically, the memory 830 can be coupled to the processor 810. The coupling in the embodiment of the present application is an indirect coupling or communication connection between devices, units or modules, which can be electrical, mechanical or other forms, and is used for information exchange between devices, units or modules. Alternatively, the processor 810 can operate in conjunction with the memory 830. The memory 830 can be a non-volatile memory, such as a hard disk drive (HDD), or a volatile memory (volatile memory), such as a random-access memory (RAM). The memory 830 is any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this.
[0176] Based on the above embodiments, embodiments of the present application further provide a computer storage medium. This storage medium stores a software program that, when read and executed by one or more processors, can implement the methods provided by any one or more of the above embodiments. The computer storage medium may include various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory, a random access memory, a magnetic disk, or an optical disk.
[0177] Based on the above embodiments, embodiments of the present application further provide a chip. The chip includes a processor configured to implement the functions described in any one or more of the above embodiments, such as acquiring or processing the data frames described in the above methods. Optionally, the chip also includes a memory configured to store the necessary program instructions and data for execution by the processor. The chip may be comprised of a single chip or may include a chip and other discrete components.
[0178] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0179] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0180] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0181] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0182] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present application without departing from the scope of the embodiments of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalents, the present application also intends to include these modifications and variations.
Claims
1. A method for transmitting data, characterized in that: Executed in a cluster system, the cluster system comprising: N service boards, N interface board sets, and M central switching boards, where N is greater than or equal to 3 and M is greater than or equal to 2; Each interface board set includes at least one interface board, the N service boards correspond to the N interface board sets one-to-one, and each service board sends or receives service data through the interface board in the corresponding interface board set; Each central switching board is in communication with at least three interface boards, and the at least three interface boards belong to different interface board sets; In response to a failure of a first interface board connected to a first central switching board among the M central switching boards, controlling the first service board to transmit first service data to the second service board based on a first strategy, and controlling the third service board to transmit second service data based on a second strategy; wherein, The first interface board belongs to the interface board set corresponding to the first business board or the second business board, the first strategy is used to indicate the first target central switching board set of the first business data, and the second strategy is used to indicate the second target central switching board set of the second business data. The first target central switching board set includes central switching boards that are connected to at least one remaining interface board in the interface board set to which the first interface board belongs except the first interface board, and the second target central switching board set includes the first central switching board.
2. The method according to claim 1, characterized in that The cluster system further includes: N cross-board sets, each of the N cross-board sets corresponding to the N service boards, each cross-board set including at least one cross-board, and each service board transmitting service data through the corresponding cross-board set and the corresponding interface board set; A first cross-connect board in the cross-connect board set corresponding to the first service board is used to transmit the first service data, and the first cross-connect board is configured with at least one logical port; In response to a failure of the first interface board, the first cross-connect board is controlled to switch to a first logical port among the at least one logical port to execute the first policy.
3. The method according to claim 1 or 2, characterized in that In response to a failure of the first interface board, it is detected that the total bandwidth of the business data that needs to be transmitted through the first target central switching board set exceeds a first threshold, and the business data in the first business data that exceeds the first threshold is suspended, where the first threshold includes the total bandwidth upper limit of the first target central switching board set.
4. The method according to any one of claims 1 to 3, characterized in that Also includes: receiving first information, where the first information is used to indicate a fault of the first interface board; According to the first information, second information is sent, where the second information is used to instruct the first service board to transmit the first service data based on the first policy.
5. The method according to claim 4, characterized in that The second information is further used to instruct the first cross-connect board to switch to the first logical port.
6. A cluster system, characterized in that: include: A control unit, N service boards, a set of N interface boards, and M central switching boards, where N is greater than or equal to 3 and M is greater than or equal to 2; Each interface board set includes at least one interface board, the N service boards correspond to the N interface board sets one-to-one, and each service board sends or receives service data through the interface board in the corresponding interface board set; Each central switching board is in communication with at least three interface boards, and the at least three interface boards belong to different interface board sets; In response to a failure of a first interface board connected to a first central switching board among the M central switching boards, the control unit controls the first service board to transmit first service data to the second service board based on a first strategy, and controls the third service board to transmit second service data based on a second strategy; wherein, The first interface board belongs to the interface board set corresponding to the first business board or the second business board, the first strategy is used to indicate the first target central switching board set of the first business data, and the second strategy is used to indicate the second target central switching board set of the second business data. The first target central switching board set includes central switching boards that are connected to at least one remaining interface board in the interface board set to which the first interface board belongs except the first interface board, and the second target central switching board set includes the first central switching board.
7. The system according to claim 6, characterized in that Also includes: N cross-board sets, each of the N cross-board sets corresponding to the N service boards, each cross-board set including at least one cross-board, and each service board transmitting service data through the corresponding cross-board set and the corresponding interface board set; A first cross-connect board in the cross-connect board set corresponding to the first service board is used to transmit the first service data, and the first cross-connect board is configured with at least one logical port; In response to a failure of the first interface board, the control unit controls the first cross-connect board to switch to a first logical port of the at least one logical port to execute the first policy.
8. The system according to claim 6 or 7, characterized in that It also includes a bandwidth detection module, In response to the failure of the first interface board, the bandwidth detection module detects that the total bandwidth of service data that needs to be transmitted through the first target central switching board set exceeds a first threshold; The control unit suspends service data in the first service data that exceeds a first threshold, where the first threshold includes a total bandwidth upper limit of the first target central switching board set.
9. The system according to any one of claims 6 to 8, characterized in that The control unit receives first information, where the first information is used to indicate a fault of the first interface board; The control unit sends second information according to the first information, where the second information is used to instruct the first service board to transmit the first service data based on the first policy.
10. The system according to claim 9, characterized in that The second information is used to instruct the first cross-connect board to switch to the first logical port.
11. The system according to any one of claims 6 to 10, characterized in that Also includes a fault detection module, The fault detection module inspects each interface board in the cluster system according to a first time period and sends a first message; The first information is used to indicate a fault of the first interface board.
12. The system according to any one of claims 6 to 10, characterized in that The first interface board includes a communication module, The first interface board reports first information through the communication module, where the first information is used to indicate a fault of the first interface board.
13. A data transmission device, characterized in that: The device comprises at least one processor configured to execute a computer program or instruction stored in at least one memory, so that the data transmission device performs the method according to any one of claims 1 to 5.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions. When the computer instructions are executed on a computer, the method according to any one of claims 1 to 5 is executed.
15. A computer program product, characterized in that The present invention comprises computer executable codes or computer executable instructions, and when the computer executable codes or computer executable instructions are executed by a computer, the method according to any one of claims 1 to 5 is implemented.