Fast scalable connector for network connection
By employing a two-stage scheduler design and dynamic adjustment of the feedback learner, the problem of rapid and scalable connectivity after large-scale network node disconnections is solved, achieving efficient network recovery and resource management, and supporting the stable operation of large-scale cluster systems.
Patent Information
- Application Number
- CN202410598550.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies struggle to quickly and scalably reconnect after large-scale network node outages, leading to performance bottlenecks and resource overruns, failing to meet the requirements of high availability and clustered systems.
A two-stage scheduler design is adopted. The main scheduler is responsible for global task management and priority sorting, while the worker scheduler executes tasks concurrently. Resources are dynamically adjusted through a feedback learner to achieve elastic scaling and optimized connection strategies.
It achieves fast and reliable network recovery, reduces service interruptions, supports instant connection of large-scale nodes and efficient resource utilization, and has fault-tolerant recovery capabilities.
Smart Images

Figure CN120956800A_ABST
Abstract
Description
Technical Field
[0001] The embodiments relate to the field of large-scale networks, and more specifically, to systems for fast and efficient connectivity of network nodes. background
[0002] In large-scale computer deployments, within any short period of time, many running network nodes (devices or servers) may go offline and then recover for various reasons (such as software or firewall upgrades, security patches, routine maintenance, power outages, etc.). In this situation, local client nodes (such as the SMx network control plane) need to be able to proactively connect to the node that first went offline and then recovered in the most efficient and fastest way to perform any subsequent services and operations.
[0003] One current approach to this situation is to use a single thread to sequentially execute reconnection tasks for lost and recovered remote nodes. This solution is simple and easy to implement, but suffers from severe performance and scalability drawbacks due to bottlenecks, especially under large-scale reconnection requirements. Another current approach is to use multiple threads independently and repeatedly execute a large number of reconnection tasks in parallel. This can provide good performance through pure parallelization, but typically lacks advanced features and other considerations (i.e., resource overruns, connection spikes, and coordination). Similarly, the timing-wheel algorithm for interleaving node reconnection may be suitable as a pure data algorithm for the underlying implementation, but does not provide a complete product solution. In general, these existing methods do not provide a suitable end-to-end approach to meet the product-readiness requirements of large-scale, high-availability, clustered system deployments.
[0004] Therefore, what is needed is a fast and scalable connector for network connectivity that minimizes communication interruptions and provides guaranteed continuous service and managed availability for thousands of remote reconnection nodes that require real-time proactive reconnection from client-peer systems.
[0005] The topics discussed in the background section should not be considered prior art simply because they are mentioned therein. Similarly, problems mentioned in or related to the topics in the background section should not be considered as having been previously recognized in the prior art. The topics in the background section merely represent different methods, which themselves can be embodiments of the invention. AXOS and AXOSDPx are trademarks of Calix Corporation. Brief description of the attached diagram
[0006] In the following figures, similar reference numerals denote similar structural elements. Although the figures depict various examples, one or more embodiments and implementations described herein are not limited to the examples depicted in the figures.
[0007] Figure 1 A system for implementing fast, scalable connectors is shown in some embodiments.
[0008] Figure 2 More detailed illustrations are shown in some embodiments. Figure 1 Fast, scalable connectors.
[0009] Figure 3 This is a block diagram illustrating the components and signal flow of a fast, scalable connector in some embodiments.
[0010] Figure 4 A group of devices that are not connected for a long time (LLnC) are shown in some embodiments.
[0011] Figure 5 This is a flowchart illustrating the overall sequence of workflows between components in a fast, scalable connector under some embodiments.
[0012] Figure 6 This is a flowchart illustrating a sequence of workflows between components in a fast, scalable connector for burst and scaling processing in some embodiments.
[0013] Figure 7 This is a flowchart illustrating a sequence of workflows between components in a fast, scalable connector for handling abnormal LLLnC nodes in some embodiments.
[0014] Figure 8 This is a flowchart illustrating the overall process of sending reconnection requests for a large number of disconnected nodes using a fast, scalable connector in some embodiments. Detailed description
[0015] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the disclosed exemplary embodiments. However, those skilled in the art will understand that the principles of the exemplary embodiments may be implemented without every specific detail. Well-known methods, processes, and components are not described in detail so as not to obscure the principles of the exemplary embodiments. Unless explicitly stated otherwise, the exemplary methods and processes described herein are not limited to a particular order or sequence, nor to a particular system configuration. Furthermore, some of the described embodiments or elements thereof may be combined, occur, or performed simultaneously, at the same time, or concurrently.
[0016] It should be noted that the described embodiments can be implemented in a variety of ways, including as a process, apparatus, system, device, method, or computer-readable medium containing computer-readable instructions or computer program code, or as a computer program product containing computer-readable program code. In the context of this disclosure, a computer-usable medium or computer-readable medium can be any physical medium that can contain or store programs used by or in conjunction with an instruction execution system, apparatus, or device.
[0017] Reference will now be made in detail to the disclosed embodiments, examples of which are illustrated in the accompanying drawings. Unless explicitly stated otherwise, the terms "transmission" and "reception" as used herein are to be understood in a broad sense, including transmission or reception in response to a particular request or without such a particular request. Thus, these terms encompass both active and passive forms of transmission and reception.
[0018] The embodiment relates to a network management system (NMS) connector that is both fast, as it can quickly establish or reconnect as soon as a remote node is ready, and scalable, as it can linearly scale to meet connectivity needs when a large number of nodes go offline and then come back online simultaneously.
[0019] Generally, connectors are software components that reliably and instantly create and manage link connections between management control plane and network device data plane nodes. An example connector is the AXOSDPx connector from Calix, which enables cable operators to deploy software-defined networking (SDN) capabilities in their access networks without disrupting their existing back-office environments. The software-based DPx connector acts as a translation layer between the back-office systems and the software-defined access operating system.
[0020] In this embodiment, the connector is designed and configured to provide rapid network recovery after a large-scale disconnection event involving numerous network nodes. A central database (DB) maintains and updates records of the current connectivity status of all devices, and a control plane consisting of multiple pairs of master schedulers and worker schedulers efficiently reconnects disconnected devices. For each pair of master and worker schedulers, the master scheduler periodically filters out a list of disconnected devices based on predefined policies (e.g., local cluster membership) and / or feedback collected by a feedback learner, and submits this list to the worker scheduler. The worker scheduler comprises two layers responsible for executing network connections. The worker scheduler is resilient, allowing the master scheduler to expand its capacity if the number of disconnected devices becomes very large, enabling the worker scheduler to concurrently connect more devices. Once the peak task period has passed, the master scheduler can reduce the capacity of the worker scheduler to free up hardware resources.
[0021] The system also features a two-tiered scheduler: an L1 scheduler and an L2 scheduler. Tasks submitted by the master scheduler first go to the L1 scheduler, and if overloaded, the L1 scheduler will load a portion of the tasks, with random delays, onto the L2 scheduler. The L2 scheduler also handles connection failures, ensuring that all devices that failed to connect after the first attempt are loaded into the L2 scheduler. In this way, devices such as those that have been inactive for a long time are de-prioritized to avoid consuming L1 scheduler resources, potentially allowing priority to be given to other devices and enabling efficient and rapid network recovery.
[0022] The feedback learner can collect key performance indicators (KPIs) in real time from the L1 and L2 schedulers and detect devices that have been disconnected for a long time. The main scheduler periodically obtains this information from the feedback learner and schedules tasks based on this component.
[0023] Figure 1 A system for implementing fast, scalable connectors at a high level is illustrated in some embodiments. For example... Figure 1 As shown, the connector system 100 consists of several components, including a data plane 104 with multiple remote nodes 108 and a control plane 102. The system 100 also has a central database 106 that persistently stores all managed network device data plane nodes 108 and provides a global connectivity status view for cluster and system recovery.
[0024] As shown, the control plane 102 includes multiple members, denoted as member 1 to member n. Each member has a set of master scheduler 110 and worker scheduler 112, which are connected to the corresponding node in the data plane 104.
[0025] like Figure 1As shown and described in more detail below, system 100 includes a control and feedback loop between a master scheduler 110 and a worker scheduler 112, as well as connection signals between a central database 106 and the master-worker scheduler and data plane 104. To provide guaranteed speed and scalability, system 100 includes a two-stage dual scheduler design. This scheduler design involves the master scheduler 110 globally managing and dispatching tasks, while the worker scheduler 112 concurrently executes these tasks for nodes 108.
[0026] Figure 2 More detailed illustrations are shown in some embodiments. Figure 1 Fast, scalable connectors. For example... Figure 2 As shown, system 200 includes a central database 202 coupled to a master scheduler 204, which is a global connection coordinator that periodically selects, prioritizes, and dispatches reconnection tasks, and is responsible for monitoring, governance, and cluster awareness. A multi-tier worker scheduler 206 is a resilient connection worker that concurrently (in parallel) executes the actual connection tasks for all disconnected nodes.
[0027] The central database 202 stores device information for all remote nodes 208, including the node managing the device, connection status, last connection time, last disconnection time, etc. Remote node 208 typically represents one or more device nodes (usually thousands) managed by the management system and accepting connection requests from the system.
[0028] The master scheduler 204 is responsible for filtering and prioritizing connection requests based on information provided by the database 202 and the feedback learner 210. It is also responsible for submitting requests to the worker scheduler 206 and scaling up / down resources within the worker scheduler. When reading initial information from the database, the master scheduler can perform filtering based on certain fields of the device information, such as manageable flags indicating whether a device should be managed by the network manager, and pre-previsioning flags indicating that a device is not yet online for management. Connection requests submitted by the master scheduler carry all the information required to establish a connection with the managed device, including the device name, device IP address, and port.
[0029] for Figure 2In this embodiment, the layered and load-balancing worker scheduler 206 comprises two layers (L1, L2) that execute locally with better isolation to distribute task traffic to the next layer during overload, thereby maximizing throughput. The worker scheduler 206 is designed to enable connectors to operate with a small resource footprint using bounded queues and a fixed-size thread pool, and to achieve automatic, on-demand vertical scaling as the workload of connection tasks surges and diminishes in a dynamic network comprising a large number of remote nodes 208.
[0030] The hierarchical worker scheduler 206 is responsible for handling connection requests from the master scheduler 204. When a request arrives within the current processing capacity of the L1 scheduler, the worker scheduler 206 will process the request directly. If the worker scheduler 206 cannot process the request in L1, it will submit the request to the L2 scheduler. The L2 scheduler uses a scalable queue and thread pool to run requests in a scheduled manner; that is, each request will be scheduled to execute at a future time. Resource consumption can vary significantly based on the number of pending requests. The maximum size of the L2 scheduler thread pool is limited by factors such as the underlying OS type, OS release, and physical memory size. Systems are typically configured to keep the thread pool size within a reasonable range to avoid excessive resource consumption and potential performance issues.
[0031] System 200 implements socket-based, I / O-layer-driven reactive rescheduling to ensure faster reconnection and minimize service connection interruptions. Connection socket I / O is reactively triggered to reschedule connection attempts based on configurable settings.
[0032] It should be noted that the terms "connection" and "reconnection" are used interchangeably to refer to effective functional coupling between components. Typically, "connection" may mean the initial connection, while "reconnection" may mean a subsequent connection after the initial connection has been interrupted. Whether connected or reconnected, both components are considered to be connected or in a connected state.
[0033] System 200 also implements feedback and learning-driven intelligent scheduling through feedback learner component 210. This ensures that the system can implement the most appropriate reconnection strategy by utilizing historical and statistical information learned from past connections. Based on the load of the elastic worker scheduler 206, whether it is running tasks, queued tasks, etc., the feedback learner 210 dynamically expands or shrinks the thread pool size of the L2 scheduler in the worker scheduler 206. The feedback learner 210 typically collects the status of each worker scheduler, including processing and pending requests, work queue depth, LLLnC devices, etc., and then provides this information to the master scheduler 204.
[0034] The system further implements node-affinity-based cluster scheduling. It uses a node-affinity-based cluster-aware approach to simplify cluster management and autonomous scheduling for large-scale production deployments.
[0035] Database 202 implements a table-based fault-tolerant recovery scheme to ensure that connectors can continue to operate in the event of failures (such as restarts, upgrades, unexpected crashes, power outages) for high availability and fault recovery systems.
[0036] System 200 also includes a cache 212 containing a cache memory set (RS) containing all devices 208 that are in a connected or pending connection state, in order to avoid duplicate connection requests from the same device.
[0037] Figure 3 This is a block diagram illustrating the components and signal flow for a fast, scalable connector in some embodiments. In an embodiment of system 300, database 302 includes node table 304, which is a tabular data element containing relevant information about devices for remote node 301. A node table typically refers to a data element (table, list, database, text document, etc.) that provides a global view of the connection status of all devices and cluster management. It also persistently stores provider data for fault recovery purposes.
[0038] In this embodiment, each remote node 301 (which may be a client and / or a server) includes one or more devices that are connected to or disconnected from the system and from each other. These states may include: connected, disconnected (disconnected or not connected), pending connection, or failed (pending disconnection). A disconnected device that is not intended to be disconnected is a device designed to reconnect as quickly as possible via the fast, scalable connector 300 to maintain overall network functionality. Such a reconnected device will then re-establish a connected state.
[0039] In this embodiment, the devices in remote node 301 that may suffer periodic failures or disconnections are established and deployed devices, not temporary or transient devices. Such devices are referred to as long-lived devices, and when they are unintentionally disconnected, they are referred to as long-unconnected (LLnC) devices. LLnC designs typically allow for prioritized / ranking connection scheduling and minimize LLnC device interference.
[0040] Depending on factors such as equipment type, equipment criticality, downtime, or duration of disconnection, LLLnC equipment may have different rankings. Figure 4 A set of long-term unconnected (LLnC) devices is illustrated in some embodiments. Figure 400 shows a set of LLnCs along a timeline ranging from minutes to hours or even longer (e.g., days, weeks, etc.), and any suitable time scale can be used. Each LLnC in example set 402 is ranked along some scale, such as from 1 to 8 along a timeline, where the LLnC's level depends on a random amount of latency for each device, which is used by the L2 scheduler to prioritize reconnecting devices, such as from LLnC_1 with a latency of 15 minutes to LLnC_8 with a latency of 2 hours, and so on. In embodiments, the level of an LLnC device determines its reconnection priority; for the illustrated example, the priority is LLnC_1 > LLnC_2 > LLnC_3 > ... > LLnC_8.
[0041] The LLnC value essentially determines the delay imposed for reconnecting a device, which causes the device to be unavailable for this additional period. In other words, the LLnC value essentially represents the scheduling priority based on a scheduled delay time in a particular implementation.
[0042] Figure 4 Provided for illustrative purposes only, and can list and rank any number of LLC devices, with the timescale set to any appropriate range.
[0043] refer to Figure 3 As described above, connector 300 includes a two-stage dual scheduler, wherein the first stage is performed by a master scheduler that prioritizes tasks, assigns tasks, and manages scheduling globally in the network, and the second stage is performed by a worker scheduler that performs reconnection locally, reschedules tasks as needed, and collects data for analysis and feedback.
[0044] As shown in system 300, node table 304 stores the connection or disconnection status of each device in node 301 and provides a fault recovery plan for connector 300. The node table provides a database-based fault-tolerant recovery scheme. Connector 300 saves remote nodes to node table 304 as a global and persistent state. Therefore, even if the application may have restarted after a failure or interruption (e.g., upgrade, software defect, etc.), the system can continue to perform reconnection scheduling.
[0045] The fault recovery information is filtered and prioritized by the master scheduler 306. In the first phase of scheduling (Phase I), the master scheduler 306 prioritizes tasks and assigns them for submission (via a "submit" command) to the worker scheduler 308. The master scheduler 306 also performs a management function 318 that monitors and scales the tasks assigned to the worker scheduler 308. It performs this globally for all disconnected devices on node 301, and the ultimately assigned tasks result in a reconnection or connection retry operation ("connection") from I / O layer 316 to node 301 via socket I / O commands.
[0046] In the second phase (Phase II), the scheduled and submitted tasks from the master scheduler 306 are then fed into the worker scheduler 308, which contains separate L1 schedulers and L2 (overflow) schedulers. Thus, the worker scheduler comprises a multi-tiered scheduler that acts as a single executor service that balances isolation (reduced interference) and load balancing.
[0047] The L1 scheduler 310 contains a bounded queue 311. The L1 scheduler 310 uses a fixed-size thread pool to process connection requests in real time. However, its capacity is limited by the queue size, so when it is overloaded, further requests are sent to the L2 scheduler 312. To maintain independence, both the L1 and L2 schedulers have their own non-shareable queues and thread pools.
[0048] The master scheduler 306 initially submits the reconnection task to the L1 scheduler 310. If the L1 scheduler accepts the task, it is passed directly to the node via the I / O layer 316. However, if the L1 scheduler is overloaded, it will further adaptively forward the task to the L2 scheduler 312 with a random delay. This delay is set by the LLnC level of the retry device and is used for load balancing purposes so that the L2 scheduler is not overwhelmed by concurrently timed reconnection tasks from the L1 scheduler.
[0049] The worker scheduler contains an unbounded queue 313, but prioritizes (or discards) connection requests based on a defined LLnC level. For devices with LLnC 8 or higher, the L2 worker scheduler may simply discard connection requests to allocate resources to requests with higher connection priority. Appropriate rules can be defined to determine reconnection priorities within the L2 scheduler. For example, it can be configured to reschedule fast retries only for devices at the LLnC 1 level, providing additional reconnection opportunities for these devices outside of the main connection cycle. Other similar rules can also be defined depending on system configuration and requirements.
[0050] The initially scheduled (from L1) or reactively rescheduled (from L2) connection task is then sent as a "connect" command from the worker scheduler 308 to the remote node 301 via the socket I / O layer 316. This socket I / O layer-driven reactive rescheduling provides better, faster connections, and the socket I / O (SKT) will perform the reactive rescheduled reconnection task on the worker scheduler. Typically, the connection command can be a generic system command that forces or creates a connection between two components.
[0051] like Figure 3 As shown, the feedback and learning circuit 314 provides intelligent scheduling based on certain collected data. The connector collects, marks, and monitors the system's performance (e.g., active threads, queued tasks, etc.) and tasks (e.g., total reconnection, scheduling delays, etc.), which can be provided in the form of statistics, historical data, trend data, expert knowledge bases, etc., to apply the most suitable and efficient reconnection strategy for a set of disconnection scenarios.
[0052] Figure 5 This is a flowchart illustrating the overall sequence of workflows between components in a fast, scalable connector under some embodiments. For example... Figure 5 As shown in Figure 500, the process flow between the master scheduler 502, the database 504, the feedback learner 506, the worker scheduler 508, and the output stage 510, which includes the I / O layer and remote nodes, is illustrated.
[0053] Database 504 provides node table 512, which specifies the endpoint, status, and cluster membership for each device of the remote nodes. Master scheduler 502 performs periodic master task processing (step 1) and accesses node table 512 to identify and select any disconnected nodes (step 2). Master scheduler 502 selects nodes managed by local cluster members (step 3) to provide device reconnection through node affinity-based cluster scheduling. To this end, the master scheduler is able to implement cluster-aware deployment and perform per-node cluster affinity scheduling to autonomously reconnect the corresponding disconnected remote nodes.
[0054] Feedback learner component 506 collects data and information from worker scheduler 508, the I / O layer, and remote node 510 to gain insights for generating feedback-driven intelligent scheduling. The master scheduler 502 then uses this insight to schedule reconnection tasks (step 5). If necessary, the master scheduler scales the worker scheduler as needed (step 6) and submits the scheduled tasks to worker scheduler 508 (step 7).
[0055] Then, the worker scheduler executes the connection task by sending a "connect" command to the remote node through the I / O layer. Output stage 510 sends an I / O callback indicating whether the connection was successful or failed back to database 504 (step 9). If necessary, such as if a previous connection attempt failed, the worker scheduler performs reactive rescheduling of the reconnection task (step 10). This reactive rescheduling is based on reactive scheduling being performed as frequently as needed (step 11). During the scheduling and rescheduling of reconnection requests, worker scheduler 508 continues to collect and provide relevant data as input to feedback learner 506 (step 12). In this way, the connector applies the reactive rescheduling process in a fine-grained manner based on the analysis of multiple information points (such as anomaly filtering, retry throttling, random delay, LLLnC matching, etc.).
[0056] like Figure 5 As shown in the processing flow, the connector utilizes a resource-efficient and flexible worker scheduler that initiates during system startup when hardware resource conditions are detected. For flexible capacity based on runtime measurements, the connector can vertically scale during bursts and decreases in reconnection tasks.
[0057] Figure 6 This is a flowchart illustrating a sequence of workflows between components in a fast, scalable connector for burst and expansion processing in some embodiments. For example... Figure 6 As shown in Figure 600, the process flow between the master scheduler 602, the database 604, the feedback learner 606, the worker scheduler 608 including L1 and L2 schedulers, and the output stage 610 including the I / O layer and remote nodes is illustrated.
[0058] In this embodiment, the feedback learner 606 acts as a collector of key performance indicators (KPIs) to dynamically collect runtime workload data from the worker scheduler 608 (L1 and L2 schedulers) (step 1). In this embodiment, the KPIs may be timed tasks or I / O event-driven and may include scheduler workload statistics, LLLnC group details, worker thread counts, queue depths, etc.
[0059] The feedback learner then gains insights about the worker scheduler (step 2). These insights can include any relevant information about device status and load. In this embodiment, this information may be provided by an operating system (OS) feature or location (such as " / Sys / WorkerScheduler / Socket / Device / ") or a similar resource.
[0060] for Figure 6In this embodiment, it is assumed that a large-scale surge event (step 3) is reported to and stored in database 604. This may be due to multiple nodes unintentionally becoming disconnected at the same time or nearly simultaneously, resulting in a large number of pending tasks (reconnection) requiring scheduling. In response, the master scheduler 602 then obtains worker scheduler workload KPI data from the feedback learner 606 to achieve feedback-driven intelligent scheduling (step 4).
[0061] Then, the master scheduler 602 determines the expansion of the worker scheduler based on task and KPI data, so as to expand the worker scheduler 608 according to the real-time demand of tasks and actual and potential scheduling overload (step 5). Expanding the worker scheduler increases its capacity as needed for better burst handling (step 6.a). Then, the L2 scheduler of worker scheduler 608 is expanded as needed (step 7). In embodiments, expansion is performed vertically by adding more virtual threads or OS threads within a single machine, and / or horizontally by adding more machines to the cluster. Other expansion schemes may also be used appropriately.
[0062] After the worker scheduler is fully expanded, the master scheduler 602 then submits connection tasks to the worker scheduler 608 in parallel (step 8.a). Then, in the output stage 610, the worker scheduler performs the connection tasks through the I / O layer and the remote node (step 9). The output stage sends I / O callbacks to the worker scheduler, and then the L1 scheduler automatically and autonomously expands or shrinks (step 10). It should be noted that the L1 scheduler expands or shrinks autonomously, while the L2 scheduler is expanded or shrunk by the master scheduler.
[0063] In case of L1 scheduler overload in step 608, overflow reconnection tasks are forwarded to L2 scheduler (step 11). L2 scheduler applies appropriate latency based on the device's LLLnC level, such as... Figure 6 As shown in step 12. These tasks are then performed by output stage 610 at the appropriate time (step 13), and then output stage 610 sends I / O callbacks to the L2 scheduler.
[0064] like Figure 6 As shown, the LLnC level information is also used by the master scheduler 602 to directly submit connection tasks to the L2 scheduler for low-priority processing (step 8.a), and the L2 scheduler can then send connection requests to the I / O layer accordingly (step 14).
[0065] As described above, the worker scheduler 608 can scale up under high connection requests or shrink under low or no requests to conserve system resources. To shrink the worker scheduler, the master scheduler 602 reduces the worker scheduler capacity to use resources more efficiently (step 6.b). This reduction can be achieved by reducing the number of virtual threads or OS threads in the worker scheduler process to reclaim resources or in a similar manner.
[0066] In some cases, certain devices or nodes may experience abnormal operation. Figure 7 This is a flowchart illustrating a sequence of workflows between components in a fast, scalable connector for handling abnormal LLLnC nodes in some embodiments. For example... Figure 7 As shown in Figure 700, the process flow between the master scheduler 702, the database 704, the feedback learner 706, the worker scheduler 708 including L1 and L2 schedulers, and the output stage 710 including the I / O layer and remote nodes is illustrated.
[0067] In this embodiment, the feedback learner 706 dynamically collects connection execution results from the I / O layer of 710. The feedback learner then generates insights about node reconnection operations (step 2).
[0068] The master scheduler 702 sends connection requests according to the LLnC level (step 3). The master scheduler 702 obtains feedback data for feedback-driven intelligent scheduling from the feedback learner 706 (step 4) and prioritizes the connection requests according to the LLnC level (which is based on the feedback learner prioritizing by the timestamp of disconnection) (step 5).
[0069] In this embodiment, the LLLnC level is a multi-range random scheduling priority used to smooth traffic surges, and the delay time is for implementation scheduling.
[0070] For high-priority scheduling, the master scheduler 702 immediately submits the non-LLnC tasks to the L1 scheduler in groups and sets a priority for each group (step 6.a). These tasks are then sent as connection tasks to the I / O layer and remote node 710 via the worker scheduler 708 (step 6.a.1). The worker scheduler sends reactive rescheduling for rapid reconnection (step 6.a.2).
[0071] For low-priority scheduling, the master scheduler 702 submits delayed LLnC tasks to the L2 scheduler in groups and sets a priority for each group (step 6.b). These tasks are then sent as selective connection tasks to the I / O layer and remote node 710 via the worker scheduler 708 (step 6.b.1). If LLnC exceeds a certain threshold (e.g., LLnC_8), the worker scheduler selectively executes socket connections and delays the execution of LLnC tasks (step 6.b.2). After the connections and I / O callbacks from the I / O layer, the worker scheduler then performs a non-reactive rescheduling of the LLnC tasks (step 6.a.3).
[0072] Figure 8 This is a flowchart illustrating the overall process of sending reconnection requests for a large number of disconnected nodes using a fast, scalable connector in some embodiments. Figure 8 The process 800 begins by receiving information from the system database about nodes that have experienced unexpected disconnections, which typically occurs on a large scale (e.g., tens of thousands to hundreds of thousands of nodes), 802.
[0073] In the first phase, the master scheduler schedules reconnection requests in parallel, which may take into account the device's LLLnC priority, 804. The master scheduler operates through its worker schedulers, which can be scaled up or down based on event demands and system configuration, 805. After scaling, the master scheduler sends the reconnection requests to the first-level (L1) scheduler of the worker schedulers, 806. The L1 scheduler has a bounded queue, so if the L1 scheduler can handle the request itself, as determined in decision step 808, it sends the request to the node, such as by using the I / O socket layer, 811. However, if the L1 scheduler is overloaded, the second-level (L2) scheduler with an unbounded queue is then activated, 810, and the request is then sent to the node, 812, or dropped if necessary.
[0074] Throughout process 800, the feedback learner collects connection, reconnection, and node state information and sends it to the master scheduler and worker scheduler, 814. The feedback learner provides insights for influencing scaling 805 and scheduling 804 steps.
[0075] In an embodiment, certain low or high priority levels can be set to determine or modify the scheduling of the master scheduler and / or worker scheduler, 816.
[0076] After the reconnection task is completed, the updated system and node status is sent to the database, 818.
[0077] The connector described in this paper provides fast and scalable network connectivity for reliable and instantaneous connections, with resource efficiency and resilient capacity. It also supports clustered deployments, fault-tolerant crash recovery, and learning-driven autonomous management. In the event of hundreds or thousands of remote nodes going offline and recovering within a short period, the implementation provides connectivity resilience with minimal service disruption, guaranteeing real-time network connectivity in independent and clustered application deployments of varying sizes.
[0078] As described above, the described functions can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, these functions can be stored on or transmitted via a computer-readable medium as one or more instructions or code, and executed by a hardware-based processing unit. A computer-readable medium can include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium. In this way, a computer-readable medium can generally correspond to a non-transitory tangible computer-readable storage medium. A data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described herein. Computer program products can include computer-readable media.
[0079] By way of example and not limitation, computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium capable of storing required program code in the form of instructions or data structures and accessible by a computer. It should be understood that computer-readable storage media and data storage media do not include carrier waves, signals, or other transient media, but rather refer to non-transient tangible storage media. The terms disks and optical discs used herein include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray™ discs, where disks typically reproduce data in a magnetized manner, and optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0080] Instructions can be executed by one or more processors (such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuits). Accordingly, the terms "processor" or "controller" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Similarly, these techniques can be implemented entirely within one or more circuit or logic elements.
[0081] The techniques disclosed herein can be implemented in various devices or apparatuses including an IC or a group of ICs (e.g., a chipset). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, various units can be combined in a hardware unit or provided by a collection of interoperable hardware units, including one or more processors as described above, and suitable software / firmware.
[0082] While one or more implementations have been described by way of example and in relation to specific embodiments, it should be understood that one or more implementations are not limited to the disclosed embodiments. Rather, it is intended to cover various modifications and similar arrangements that would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation to include all such modifications and similar arrangements.
Claims
1. A method for reconnecting a device to a network after an unexpected disconnection, comprising: Receive information from the database about nodes that experienced unexpected disconnections; The master scheduler concurrently schedules reconnection requests to be sent to the node; The request is sent to the first-level scheduler of the worker scheduler; If the first-layer scheduler has sufficient resource capacity to process the request, then the request is sent from the first-layer scheduler to the node; and If the first-level scheduler does not have sufficient resource capacity, then the second-level scheduler of the worker scheduler sends redundant requests to the node.
2. The method according to claim 1, wherein, The disconnection includes large-scale system outages involving thousands of nodes.
3. The method according to claim 1, wherein, The first-level scheduler includes a bounded queue for storing the requests, and the second-level scheduler includes an unbounded queue for processing the excess requests.
4. The method of claim 1, further comprising sending the reconnection request to the node using a socket-based input / output (I / O) layer.
5. The method of claim 1, further comprising defining a low or high priority level for each of the nodes, wherein, The priority level determines the priority for scheduling reconnection requests for the corresponding node.
6. The method of claim 5 further includes designating low-priority nodes as long-term unconnected (LLnC) nodes.
7. The method of claim 6, further comprising allocating a random time delay to the LLnC node to delay the time for scheduling the reconnection request for the LLnC node, and wherein, The random time delay is selected from a range of possible time delay values ranging from a few minutes to several hours.
8. The method of claim 1, further comprising updating the database with reconnection information after the node executes the reconnection request.
9. The method according to claim 1, wherein, The master scheduler and worker scheduler are maintained in a control plane coupled to the database, and the nodes are maintained in a data plane coupled to the control plane.
10. The method according to claim 1, further comprising: The worker scheduler is scaled based on system configuration, request volume, and feedback information to accommodate the reconnection requests; and The feedback learner collects node and connection information to provide the feedback information.
11. A system for reconnecting a device to a network after an unexpected disconnection, comprising: A database that receives information about nodes that have experienced unexpected disconnections; The master scheduler concurrently schedules reconnection requests to be sent to the node; and A worker scheduler having a first-level scheduler that receives the request from the master scheduler, wherein if the first-level scheduler has sufficient resource capacity to process the request, the first-level scheduler sends the request to the node; otherwise, the first-level scheduler sends any excess requests to a second-level scheduler for transmission to the node.
12. The system according to claim 11, wherein, The master scheduler is scaled to accommodate the reconnection requests based on system configuration, request volume, and feedback information.
13. The system of claim 12 further includes a feedback learner that collects node and connection information to provide the feedback information.
14. The system according to claim 11, wherein, The first-level scheduler includes a bounded queue for storing the requests, and the second-level scheduler includes an unbounded queue for processing the excess requests.
15. The system of claim 14, further comprising a socket-based I / O layer that sends the reconnection request to the node.
16. The system according to claim 11, wherein, The node is defined as having a low or high priority level, and further wherein the priority level determines the priority for scheduling reconnection requests for the corresponding node, and further wherein the low priority level node is designated as a long-term inactive (LLnC) node, and further wherein the LLnC node is assigned a random time delay to delay the time for scheduling the reconnection request for the LLnC node.
17. The system according to claim 11, wherein, The master scheduler and worker scheduler are maintained in a control plane coupled to the database, and the nodes are maintained in a data plane coupled to the control plane.
18. A system for reconnecting a device to a network after an unexpected disconnection, comprising: A central database that receives and stores information about the status and connectivity of nodes that have experienced unexpected disconnections; A data plane, which contains the nodes; and The control plane maintains a master scheduler and a multi-layered scalable worker scheduler, wherein the master scheduler prioritizes reconnection requests and assigns them to the data plane during a first reconnection phase, and the worker scheduler performs reconnection tasks locally and collects statistics and data from a feedback learner during a second reconnection phase to modify the scaling of the worker scheduler and the prioritization of the reconnection requests.
19. The system according to claim 18, wherein, The worker scheduler is scaled by the master scheduler to accommodate the reconnection requests based on system configuration, request volume, and statistics and data from the feedback learner.
20. The system according to claim 19, wherein, The worker scheduler includes a first-level scheduler that receives the request from the master scheduler, wherein if the first-level scheduler has sufficient resource capacity to process the request, the first-level scheduler sends the request to the node; otherwise, the first-level scheduler sends any excess requests to a second-level scheduler for transmission to the node.