Intelligent optical fiber distribution cooperative scheduling method, device, equipment and medium

Through the three-layer collaborative architecture of federated learning and reinforcement learning, the health of fiber-optic wiring robots is evaluated and resources are dynamically scheduled, which solves the problems of traditional low operation and maintenance efficiency and insufficient multi-machine collaboration capabilities, and realizes intelligent fiber network management.

CN120263715AActive Publication Date: 2025-07-04BEIJING RUIQI HAODI TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510748484.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Traditional manual operation and maintenance efficiency is low, the error operation rate is high, and the capacity of single-machine automation equipment is limited, making it difficult to meet the needs of large-scale networking. The existing technology lacks multi-machine collaboration capabilities and cannot adapt to the needs of rapid service activation and recovery.

Method used

The global LSTM model of federated learning training is used to evaluate the health of fiber-optic wiring robots, abstract the fiber core into a logical resource pool, and resource scheduling is carried out based on dynamic weights and reinforcement learning algorithms. A three-layer collaborative architecture is built to realize dynamic scheduling of intelligent fiber-optic wiring.

Benefits of technology

It realizes intelligent operation and maintenance management, reduces manual intervention, reduces human errors, improves operation and maintenance efficiency, adapts to dynamic changes in network status, and optimizes resource allocation and path selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263715A_ABST
    Figure CN120263715A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent optical fiber distribution cooperative scheduling method, device and equipment and a medium, and belongs to the technical field of optical fiber communication, and the method comprises the steps: training a global LSTM model based on federated learning, obtaining a global model parameter, carrying the global model parameter into a local LSTM model of each intelligent optical fiber distribution robot, and evaluating the health degree of each intelligent optical fiber distribution robot; abstracting fiber cores of different intelligent optical fiber distribution robots into a logic resource pool and calculating dynamic weights of the logic resource pool; and determining an optimal resource range based on the dynamic weight, and determining a scheduling path meeting service requirements from the optimal resource range based on a search algorithm and a reinforcement learning algorithm. Therefore, federated learning, dynamic weight scheduling and reinforcement learning are fused, a three-layer collaborative architecture is constructed, resource allocation and path selection decision can be automatically carried out according to different service requirements, manual intervention is reduced, human errors are reduced, and intelligent operation and maintenance management is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of optical fiber communication technologies, and more particularly, to an intelligent optical fiber distribution collaborative scheduling method, device, equipment, and medium. Background Art

[0002] With the development of technologies such as 5G and industrial Internet, the number of optical cable cores in computer rooms has increased sharply (up to thousands of cores in a single computer room). Traditional manual operation and maintenance has low efficiency and high error rates, and the capacity of single-machine automation equipment is limited, making it difficult to meet the needs of large-scale networking. Existing technologies urgently need to break through bottlenecks such as rigid resource scheduling and insufficient multi-machine collaboration capabilities to achieve intelligent and highly reliable dynamic operation and maintenance. Currently, optical fiber network operation and maintenance mainly rely on manual operation or single-machine automation equipment. Manual operation and maintenance requires technical personnel to plug and unplug fiber optic jumpers on-site, which is time-consuming and may have errors in manual operations, and cannot meet the usage requirements of rapid service activation and restoration; although single-machine robots can operate automatically, their capacity is limited, and they lack multi-machine collaboration capabilities. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide an intelligent optical fiber distribution collaborative scheduling method, device, equipment, and medium, which can achieve dynamic scheduling of optical fiber resources within a cluster through a three-layer collaborative architecture of federated learning prediction - dynamic weight scheduling - reinforcement learning decision-making.

[0004] An intelligent optical fiber distribution collaborative scheduling method provided by an embodiment of this application is applied to a cluster composed of multiple intelligent optical fiber distribution robots. The method includes the following steps: Based on federated learning, train a global LSTM model for predicting the health status of intelligent optical fiber distribution robots to obtain global model parameters, and load the global model parameters into the local LSTM models of each intelligent optical fiber distribution robot to evaluate the health of each intelligent optical fiber distribution robot; Abstract the cores of different intelligent optical fiber distribution robots into a logical resource pool, and calculate its dynamic weight according to the load condition of each logical resource pool and the health of the intelligent optical fiber distribution robots included; Based on the dynamic weight, determine the optimal resource range, and determine a scheduling path that meets the service requirements from the optimal resource range based on a search algorithm and a reinforcement learning algorithm.

[0005] In some embodiments, the method further includes the following steps: When it is monitored that the network, device, and link status change or the service request volume fluctuates beyond a set threshold, recalculate the dynamic weight and determine the scheduling path.

[0006] In some embodiments, training the global LSTM model based on federated learning for predicting the health status of intelligent optical fiber distribution robots to obtain global model parameters includes the following steps: Initialize the global LSTM model and send the initialized global model parameters to the local LSTM models of each intelligent optical fiber distribution robot; Preprocess the device data collected by the intelligent optical fiber distribution robots, construct time series data of different modalities, and train the local LSTM models based on the time series data; Execute parameter alignment compensation and parameter aggregation strategies on the global LSTM model using the local training parameters obtained by each intelligent optical fiber distribution robot to obtain updated global model parameters.

[0007] In some embodiments, the device data includes optical power loss, the number of pluggings and unpluggings of optical fiber connectors, the running mileage of the lead screw, and environmental data. The preprocessing of the device data collected by the intelligent optical fiber distribution robots to construct time series data of different modalities includes the following steps: Statistically calculate the number of pluggings and unpluggings within each window based on the sliding window mechanism, and divide it into low-frequency plugging and unplugging scenarios and high-frequency plugging and unplugging scenarios according to a preset threshold value; among them, in the low-frequency plugging and unplugging scenario, perform extreme value normalization processing on the optical power loss; in the high-frequency plugging and unplugging scenario, perform standard deviation normalization processing on the optical power loss; Calculate the cumulative fatigue degree based on the number of pluggings and unpluggings, the running mileage of the lead screw, and the time decay effect, which is used to measure the cumulative pressure and loss borne by the device during use; Obtain the time stamp of each plugging and unplugging operation, and convert it into a first time code and a second time code according to the periodicity of the plugging and unplugging of the intelligent optical fiber distribution robot; Align the normalized optical power loss, cumulative fatigue degree, number of pluggings and unpluggings, first time code, second time code, and environmental data according to time steps to form a two-dimensional input tensor as time series data of different modalities.

[0008] In some embodiments, determining the optimal resource range based on the dynamic weight includes the following steps: Determine the priority of the logical resource pool according to the magnitude sorting of the dynamic weights, and determine the priority of each intelligent optical fiber distribution robot in the logical resource pool according to the health degree and the number of idle fiber cores of the intelligent optical fiber distribution robot; Select a set number of logical resource pools and the intelligent optical fiber distribution robots they contain in the order of priority as the optimal resource range.

[0009] In some embodiments, determining a scheduling path that meets service requirements from the optimal resource range based on a search algorithm and a reinforcement learning algorithm includes the following steps: Using a search algorithm based on network topology information, calculate a set of feasible paths from the source node to the destination node from the optimal resource range; wherein, if there is no path that meets the service requirements, expand the selected logical resource pool and the number of intelligent fiber distribution robots it contains in the order of priority, and re-determine the optimal resource range until a path that meets the service requirements is found; Obtain the state of the current optical fiber network, and select a path from the set of feasible paths as an action to execute using a greedy strategy; wherein, if the generated random number is less than the set exploration rate, randomly select a path from the set of feasible paths as an action to execute, and if the generated random number is greater than the set exploration rate, select the path with the maximum cumulative reward expectation value from the set of feasible paths as an action to execute; Calculate the immediate reward after the action execution based on the set reward strategy, and calculate the cumulative reward expectation value after the action execution based on the immediate reward, the set learning rate and discount factor, the cumulative reward expectation value when the action is not executed, and the maximum cumulative reward expectation value among all feasible paths; Based on the calculated cumulative reward expectation value, perform iterative optimization of the path to determine a scheduling path that meets the service requirements.

[0010] In some embodiments, the search algorithm uses the A* algorithm, and the state of the optical fiber network includes network topology information, the dynamic weight of the logical resource pool, and the health of the intelligent fiber distribution robot; the network topology information includes the connection relationship between nodes, the link length, and the optical power attenuation.

[0011] In some embodiments, there is also provided an intelligent fiber distribution collaborative scheduling device, which is applied to a cluster composed of multiple intelligent fiber distribution robots. The device includes: A federated learning health prediction module, which is used to train a global LSTM model for predicting the health status of intelligent fiber distribution robots based on federated learning to obtain global model parameters, and load the global model parameters into the local LSTM models of each intelligent fiber distribution robot to evaluate the health of each intelligent fiber distribution robot; A dynamic weight partitioning module, which is used to abstract the cores of different intelligent fiber distribution robots into logical resource pools, and calculate their dynamic weights according to the load conditions of each logical resource pool and the health of the intelligent fiber distribution robots it contains; A reinforcement learning path matching module, which is used to determine the optimal resource range based on the dynamic weights, and determine a scheduling path that meets the service requirements from the optimal resource range based on a search algorithm and a reinforcement learning algorithm.

[0012] In some embodiments, an electronic device is further provided, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of any one of the above-mentioned intelligent optical fiber distribution collaborative scheduling methods are executed.

[0013] In some embodiments, a computer-readable storage medium is further provided. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of any one of the above-mentioned intelligent optical fiber distribution collaborative scheduling methods are executed.

[0014] An intelligent optical fiber distribution collaborative scheduling method, device, equipment, and medium described in this application are applied to a cluster composed of multiple intelligent optical fiber distribution robots. Based on federated learning, a global LSTM model for predicting the health status of intelligent optical fiber distribution robots is trained to obtain global model parameters, and the global model parameters are loaded into the local LSTM models of each intelligent optical fiber distribution robot to evaluate the health of each intelligent optical fiber distribution robot; the fiber cores of different intelligent optical fiber distribution robots are abstracted into a logical resource pool, and the dynamic weight of each logical resource pool is calculated according to the load condition of each logical resource pool and the health of the intelligent optical fiber distribution robots included; the optimal resource range is determined based on the dynamic weight, and a scheduling path that meets the service requirements is determined from the optimal resource range based on a search algorithm and a reinforcement learning algorithm. Thus, by integrating federated learning, dynamic weight scheduling, and reinforcement learning, a three-layer collaborative architecture is constructed, which can automatically make resource allocation and path selection decisions according to different service requirements, reduce manual intervention, reduce human errors, and achieve intelligent operation and maintenance management. Description of the Drawings

[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 Shows the flowchart of the intelligent optical fiber distribution collaborative scheduling method described in the embodiments of this application; Figure 2 Shows the flowchart of training a global LSTM model for predicting the health status of intelligent optical fiber distribution robots based on federated learning to obtain global model parameters in the embodiments of this application; Figure 3The flowchart shows the method for determining a scheduling path that meets service requirements from the optimal resource range based on a search algorithm and a reinforcement learning algorithm according to an embodiment of the present application; Figure 4 The schematic structural diagram shows an intelligent optical fiber distribution collaborative scheduling device according to an embodiment of the present application; Figure 5 The schematic structural diagram shows an electronic device according to an embodiment of the present application. Detailed implementation manners

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn to actual scale. The flowcharts used in the present application show operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.

[0018] In addition, the described embodiments are only some embodiments of the present application, rather than all embodiments. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the protection scope of the present application.

[0019] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the existence of the features stated hereinafter, but does not exclude the addition of other features.

[0020] In view of the technical problems proposed in the background art, the present application provides an intelligent optical fiber distribution collaborative scheduling method, device, equipment, and medium, which can realize the dynamic scheduling of optical fiber resources within a cluster through a three-layer collaborative architecture of federated learning prediction-dynamic weight scheduling-reinforcement learning decision-making.

[0021] See the appended specification Figure 1 A kind of intelligent optical fiber distribution collaborative scheduling method provided by the present application is applied to a cluster composed of multiple intelligent optical fiber distribution robots, and includes the following steps: S1. Train a global LSTM model for predicting the health status of intelligent optical fiber distribution robots based on federated learning to obtain global model parameters, and load the global model parameters into the local LSTM models of each intelligent optical fiber distribution robot to evaluate the health of each intelligent optical fiber distribution robot. S2. Abstract the fiber cores of different intelligent optical fiber distribution robots into a logical resource pool, and calculate its dynamic weight according to the load condition of each logical resource pool and the health of the intelligent optical fiber distribution robots included. S3. Determine the optimal resource range based on the dynamic weight, and determine the scheduling path that meets the service requirements from the optimal resource range based on the search algorithm and the reinforcement learning algorithm.

[0022] In step S1, mainly through the distributed federated learning framework, the distributed training of local data of intelligent optical fiber distribution robots and the aggregation of global models are realized to improve the adaptability and accuracy of the health monitoring of intelligent optical fiber distribution robots.

[0023] See the attached Figure 2 description. The steps of training a global LSTM model for predicting the health status of intelligent optical fiber distribution robots based on federated learning to obtain global model parameters include the following: S101. Initialize the global LSTM model, and send the initialized global model parameters to the local LSTM models of each intelligent optical fiber distribution robot. S102. Preprocess the device data collected by the intelligent optical fiber distribution robot, construct time series data of different modalities, and train the local LSTM model based on the time series data. S103. Execute the parameter alignment compensation and parameter aggregation strategy on the global LSTM model using the local training parameters obtained by each intelligent optical fiber distribution robot to obtain the updated global model parameters.

[0024] Specifically, in step S101, the constructed global LSTM model for predicting the health status of intelligent optical fiber distribution robots can effectively process time series data, capture the variation laws of connector plugging and unplugging frequency, cumulative fatigue degree, etc. over time. After the model is initialized, global model parameters such as LSTM layer weights, fully connected layer weights, LSTM layer biases, and fully connected layer biases are sent to each intelligent optical fiber distribution robot.

[0025] In step S102, after each intelligent optical fiber distribution robot receives the initialized global model parameters, it trains the local LSTM model of the intelligent optical fiber distribution robot based on the locally collected and preprocessed time series data. Among them, the device data collected by the intelligent optical fiber distribution robot includes optical, mechanical, and environmental data such as the optical power loss of the optical fiber channel, the number of insertions and removals of the optical fiber connector, temperature, humidity, vibration intensity (time series data), and the mileage of the lead screw operation. And data preprocessing is carried out in the following ways: (1) Based on the sliding window mechanism, count the number of insertions and removals within each window, and divide it into low-frequency insertion / removal scenarios and high-frequency insertion / removal scenarios according to a preset threshold value; among them, in the low-frequency insertion / removal scenario, extreme value normalization is performed on the optical power loss; in the high-frequency insertion / removal scenario, standard deviation normalization is performed on the optical power loss.

[0026] That is, dynamic normalization of the optical power loss is performed by segmenting according to the insertion / removal frequency. This is because, in the low-frequency insertion / removal scenario, there are few insertion / removal operations, the optical power loss data is relatively stable, and the noise ratio is high; in the high-frequency insertion / removal scenario, insertions and removals are frequent, and the optical power loss fluctuates violently (such as a sharp increase in loss during insertion / removal). Then, the optical power loss data is divided into multiple subsets according to the insertion / removal frequency interval (such as low and high), and normalization parameters (mean, standard deviation, or extreme value) are calculated independently for each subset, so as to adapt to the data distribution differences at different frequencies. By segmenting according to the insertion / removal frequency, "high-frequency amplification and low-frequency smoothing" are realized, directly associating the fluctuation characteristics of the optical power loss with the insertion / removal frequency, and improving the model's response ability to extreme scenarios.

[0027] Specifically, when implementing, the short-term time window Tshort (in days) can be defined through the sliding window mechanism, and the number of insertions and removals F within each window is counted. Due to the existence of low-frequency and high-frequency insertion / removal scenarios, on the basis of counting the number of insertions and removals in the time window, the counting of the number of insertions and removals in the long-term time window Tlong is added, where Tlong is multiple consecutive time windows Tshort. The average insertion / removal frequency of the Tlong time window is evaluated by the exponentially weighted moving average (EWMA).

[0028]

[0029] Among them, is the forgetting factor, the larger it is, the smaller the influence of historical data. Ft represents the number of insertions and removals in the t time window; represents the historical average number of insertions and removals before the t time window. In one embodiment, the average insertion / removal frequency of the Tlong time window is divided into low frequency and high frequency through a preset threshold value.

[0030] In the low-frequency plugging and unplugging scenario, extreme value normalization is performed on the optical power loss.

[0031] Take the maximum and minimum values within the Tlong time window, and calculate the normalized optical power loss data:

[0032] Where P is the optical power loss data collected by a certain intelligent optical fiber distribution robot, is the calculated normalized data, is the maximum value of P within the Tlong time window, is the minimum value of P within the Tlong time window, S is the scaling factor (value range: 0.7 - 0.9, default: 0.8), O is the offset (default 0.1). Setting S and O aims to compress the normalization range to the middle interval, ensure that the minimum value is not zero, and guarantee the physical meaning of the numerical value.

[0033] In the high-frequency plugging and unplugging scenario, standard deviation normalization is performed on the optical power loss.

[0034]

[0035] Where, is the mean value of the optical power loss within the window, is the standard deviation of the optical power loss within the window.

[0036] (2) Calculate the cumulative fatigue degree AEF based on the number of plugging and unplugging times, the running mileage of the lead screw, and the time decay effect, so as to measure the cumulative pressure and loss endured by the equipment during use.

[0037]

[0038] Where, is the plugging and unplugging dynamic attenuation, is the basic damage amount per single plugging and unplugging, usually 0.15; is the overload force multiple, usually 1.2, which takes effect when Fk > 0.8 Fmax; is the stress relaxation coefficient, for fiber optic connectors, it is usually 0.02 / hour according to the material; is the dynamic wear of the lead screw mileage, is the basic wear rate, usually 0.003 for stainless steel guide rails; is the non-linear wear acceleration index, usually 2.1; The speed threshold refers to that during the operation of the intelligent optical fiber distribution robot equipment, when the lead screw movement speed reaches a certain value, it will have a significant impact on the wear of the equipment, and this value is the speed threshold. Its value is determined according to the mechanical properties and design specifications of the equipment.

[0039] (3) Convert the timestamp of each plugging and unplugging operation into a time encoding.

[0040] Among them, the timestamp of the plugging and unplugging operation records the specific time when each plugging and unplugging operation occurs. It can provide information about the time pattern and periodicity of the plugging and unplugging operations. Convert the timestamp into a numerical format suitable for model processing as the input dataset for collection. Convert the plugging and unplugging timestamp into a time interval, that is, the time difference (in hours) between the current time and the most recent plugging and unplugging. According to the periodic characteristics of the plugging and unplugging of the intelligent optical fiber distribution robot, select monthly and intraday periodicities for the periodic information of the plugging and unplugging operations. Among them, time encoding 1 represents periodicity within 30 days: convert the timestamp into the number of hours within 30 days (0 - 720), and embed it through sine / cosine encoding: ; Time encoding 2 represents intraday periodicity: convert the timestamp into the number of hours within the day (0 - 23), and similarly, perform sine / cosine encoding: .

[0041] Furthermore, align the time series data of different modalities according to time steps to form a two-dimensional input tensor , as the time series data of different modalities. Among them, considering that the plugging and unplugging situation of the optical fiber connector has a greater impact on the performance and health of the intelligent optical fiber distribution robot, the number of plugging and unplugging times of the optical fiber connector at each time step is regarded as a new feature and added to Xt. The Min-Max normalization method can be used to convert the number of plugging and unplugging times n into .

[0042] Among them, when training the local LSTM model of the intelligent optical fiber distribution robot based on the locally collected and preprocessed time series data, calculate the mean square error between the predicted value and the true value through the mean square error loss function MSE to evaluate the accuracy of the model prediction.

[0043]

[0044] Among them, L represents the value of the loss function, which is used to measure the difference degree between the model predicted value and the true value. The smaller the value of the loss function, the closer the model prediction result is to the true value, and the better the performance of the model. N represents the total number of time series data samples participating in the model training, including information such as optical power loss, cumulative fatigue degree, and plugging and unplugging frequency. The more the number of samples, the more comprehensive the features learned by the model, and the stronger the generalization ability of the model training result; represents the true health of the intelligent optical fiber distribution robot labeled for the i-th sample; represents the predicted health value of the intelligent optical fiber distribution robot obtained by the model analyzing the i-th sample, that is, the model according to the currently set parameters, for the input sample data The output result obtained after performing forward propagation calculation. During the training process of the local LSTM model, by continuously adjusting the model parameters, is made as close as possible to , thereby reducing the value of the loss function L.

[0045] Furthermore, during the training process of the intelligent optical fiber distribution robot health status prediction model, considering the imbalance of data volume of different intelligent optical fiber distribution robots and the processing efficiency of intelligent optical fiber distribution robots limited by computing power, the momentum stochastic gradient descent algorithm (Momentum SGD) is used to adjust the model parameters to obtain the smallest possible L value. The parameter adjustment process is as follows:

[0046] Among them, represents the parameter value of the model at the t-th iteration, that is, the latest parameter value obtained after this update. represents the parameter value of the model at the -th iteration, that is, the parameter value at the end of the previous iteration. represents the momentum at the t-th iteration, representing the cumulative effect of the gradient update direction; is calculated by the following formula:

[0047] Among them, t represents the current iteration step, identifying the current iteration number. The model will be iterated multiple times during training, and the model parameters will be updated each time. is the momentum coefficient, and its value range is between [0, 1], usually taking the value of 0.9. determines the influence degree of the previous momentum on the current momentum. The larger μ is, the greater the influence of the past gradient information on the current update, and the more the model can maintain the previous movement direction during the update process; the smaller μ is, the relatively greater the influence of the current gradient. is the momentum at the -th iteration, representing the momentum value calculated in the previous iteration. represents the learning rate, which controls the step size of each parameter update. If the learning rate is too large, the model may oscillate or even diverge near the optimal solution; if the learning rate is too small, the convergence speed of the model will become very slow. represents the gradient of the loss function L with respect to the parameter θ at the -th iteration. The gradient represents the change rate of the loss function at the current parameter value, and its direction points to the direction where the loss function increases fastest, and the negative gradient direction is the direction where the loss function decreases fastest.

[0048] In step S103, after the intelligent fiber optic wiring robot completes the local LSTM model parameter optimization, it feeds back the updated LSTM layer weights, fully connected layer weights, LSTM layer biases and fully connected layer bias parameters to the global LSTM model, performs parameter alignment compensation and parameter aggregation strategies to obtain updated global model parameters.

[0049] Specifically, since there are large differences in the normalized parameters of the optical power of low-frequency plugging and high-frequency plugging intelligent fiber optic wiring robots, before federal aggregation, the local standard deviation and global standard deviation of the accumulated cycle statistics are used to perform feature scaling compensation:

[0050] in, The original model parameters of the i-th intelligent optical fiber wiring robot obtained after the local LSTM model training are obtained based on the local data of device i and reflect the operating characteristics of device i itself; It is the standard deviation of the local optical power loss of the i-th intelligent optical fiber distribution robot within the set accumulation period, reflecting the discrete degree of the local optical power loss data of device i. It is compared with the global standard deviation to determine the degree of difference between the data characteristics of device i and the global data characteristics; It is the standard deviation of the current optical power loss data of all intelligent fiber optic wiring robots, representing the discrete degree of global data, and is used to unify the characteristic scales of different intelligent fiber optic wiring robots. By calculating the standard deviation of the optical power loss data of all devices, a global standard can be obtained to measure the fluctuation of data of different devices and avoid the influence of low-frequency plugging and unplugging of intelligent fiber optic wiring robots on the model. After feature scaling compensation, the local model parameters of the i-th device are obtained. This parameter not only considers the local optical power loss data characteristics of device i, but also adjusts the parameters by comparing with the global standard deviation to make it more consistent with the requirements of the global model parameters.

[0051] The plugging and unplugging frequency reflects the difference in the intensity of equipment use. Different intelligent fiber optic wiring robots have different plugging and unplugging frequencies, which affects the optical power loss. When plugging and unplugging at low frequencies, the optical power loss data is relatively stable but the noise ratio is high; when plugging and unplugging at high frequencies, the optical power loss will surge and fluctuate violently at the moment of plugging and unplugging. Considering the weight influence of low-frequency plugging and unplugging devices and high-frequency plugging and unplugging devices during parameter aggregation can more accurately learn the relationship between optical power loss and plugging and unplugging operations, improve the model's ability to respond to the health status of the equipment in extreme plugging and unplugging scenarios, and ensure that the model output is more in line with the actual operation of the equipment.

[0052] During the long-term operation of intelligent optical fiber distribution robots, optical and mechanical components continuously wear and age. Among them, the cumulative fatigue degree combines factors such as the comprehensive number of plug-and-play operations and the running mileage of the lead screw, and combines with the time decay effect to measure the cumulative pressure and loss borne by the equipment. Introducing the cumulative fatigue degree enables the model to comprehensively consider the health risks accumulated during the long-term operation of the equipment, avoid only focusing on the current plug-and-play operations, supplement information from the dimension of long-term equipment aging, and cooperate with short-term performance indicators such as optical power loss.

[0053]

[0054] Among them, represents the global model parameters updated after parameter aggregation, by performing weighted summation on the local model parameters of each. is the number of local optical power loss data samples of the i-th intelligent optical fiber distribution robot; is the sum of the number of local optical power loss data samples of all intelligent optical fiber distribution robots; is the historical average plug-and-play frequency (times / hour) of the i-th intelligent optical fiber distribution robot; is the global average plug-and-play frequency, measuring the plug-and-play frequency level of the overall equipment; is the cumulative fatigue degree of the i-th intelligent optical fiber distribution robot; represents the average cumulative fatigue degree of all intelligent optical fiber distribution robots, measuring the fatigue degree level of the overall equipment; is the adjustment factor, increasing it will enhance the weight of intelligent optical fiber distribution robots with high-frequency plug-and-play, usually being 0.5; is the activity coefficient, determined according to the number of plug-and-play operations in the past 30 days (such as =1 + the number of plug-and-play operations in the past 30 days), so as to ensure that intelligent optical fiber distribution robots with different plug-and-play frequencies contribute more reasonably to the calculation results of the global model parameters.

[0055] By considering the differences in data volume, plug-and-play frequency, and activity of different devices in a weighted manner, for intelligent optical fiber distribution robots with more plug-and-play operations, their parameters have a greater weight during the aggregation of the global model. At the same time, the characteristics of intelligent optical fiber distribution robots with low-frequency plug-and-play are not ignored, making the global model integrate the information of each local model more reasonably, and improving the adaptability and prediction accuracy for the health status of different devices.

[0056] After the global LSTM model completes double weight aggregation and parameter alignment, the updated global model parameters Send them to each intelligent optical fiber distribution robot. Each intelligent optical fiber distribution robot collects optical fiber-related data at the current moment and for a certain period of historical time (such as the past 24 hours), including optical power loss, plugging and unplugging times, temperature, humidity, etc. Organize these data into time series data in the same format as during training according to time steps, construct the input tensor X, and input the prepared input tensor X into the trained local LSTM model. The model outputs the predicted health value of the intelligent optical fiber distribution robot.

[0057] Based on the federated learning framework, it solves the problems that it is difficult for different intelligent optical fiber distribution robot models to be collaboratively optimized and unable to fully utilize the overall data characteristics. Through distributed modeling and global model aggregation, it can share the training results of each intelligent optical fiber distribution robot and improve the model performance and adaptability. This enables the model to more comprehensively capture the operation rules of each intelligent optical fiber distribution robot in the network, accurately evaluate the health status of the equipment, meet the operation and maintenance needs of different intelligent optical fiber distribution robots in large-scale optical fiber networks, and improve the overall operation and maintenance efficiency of the optical fiber network. And embed plugging and unplugging event markers and fatigue accumulation counters in the LSTM model. This method solves the problem of difficult to capture the combined effects of instantaneous operations and long-term aging of intelligent optical fiber distribution robots, breaks through the limitation that traditional models can only process single-type data or static data, can dynamically and comprehensively model the equipment state, provides a guarantee for accurately evaluating the equipment health, and then timely discovers potential problems of the equipment, ensures the stable operation of intelligent optical fiber distribution robots, and enhances the reliability of the optical fiber network. In addition, consider plugging and unplugging frequency and cumulative fatigue during parameter aggregation for dual weight allocation, and at the same time execute the event alignment compensation mechanism. This solves the generalization problem of multi-device heterogeneous data, improves the limitation that traditional federated aggregation strategies are difficult to handle data differences between devices, ensures that the global model reasonably integrates the local model information of each intelligent optical fiber distribution robot, enhances the adaptability and prediction accuracy for the health status of different devices, and enables the model to be better applied to actual complex intelligent optical fiber distribution robot scenarios.

[0058] In step S2, mainly construct a logical resource pool. Specifically, abstract the optical cable cores connected to different intelligent optical fiber distribution robots into a logical resource pool. Assign a unique logical identifier to the core range managed by each intelligent optical fiber distribution robot. For example, logical resource pool 1 includes two intelligent optical fiber distribution robots A and B. Intelligent optical fiber distribution robot A manages cores numbered 1 - 288; intelligent optical fiber distribution robot B manages cores numbered 289 - 480, and the logical identifier of the resource pool corresponding to intelligent optical fiber distribution robots A and B is Pool_1; logical resource pool 2 includes one intelligent optical fiber distribution robot C, intelligent optical fiber distribution robot C manages cores numbered 481 - 768, and the corresponding logical identifier of the resource pool is Pool_2. Then integrate the logical identifiers of the cores managed by all intelligent optical fiber distribution robots to form a global logical resource pool.

[0059] When a service request requires selecting different fiber cores connected by two intelligent optical fiber distribution robots, calculate their dynamic weights according to the real-time load conditions of each logical resource pool and the average health of the fiber cores , so as to balance the load of the intelligent optical fiber distribution robots and preferentially assign tasks to the fiber core areas managed by the intelligent optical fiber distribution robots with high health.

[0060]

[0061] Among them, is the dynamic weight of the i-th logical resource pool. The higher the weight, the more preferentially the logical resource pool is selected during task allocation; , are the weight coefficients used to adjust the relative importance of the load conditions and the average health of the fiber cores, ; is the number of idle fiber cores in the i-th logical resource pool, reflecting the load condition of the logical resource pool. The more idle fiber cores, the lower the load; is the total number of fiber cores in the i-th logical resource pool; is the average health of the fiber cores in the i-th logical resource pool, which is obtained by averaging the health scores of all fiber cores in the logical resource pool, reflecting the health status of the fiber cores. The higher the score, the better the health of the fiber cores.

[0062]

[0063] Among them, is the number of intelligent optical fiber distribution robots in the logical resource pool; represents the health score of the i-th intelligent optical fiber distribution robot with a value range of 0-1, is the health prediction value of the i-th intelligent optical fiber distribution robot, expressed in percentage; is the number of idle fiber cores connected by the i-th intelligent optical fiber distribution robot. This parameter reflects the remaining resources available for task allocation by the intelligent optical fiber distribution robot. The more idle fiber cores, the more the intelligent optical fiber distribution robot should be considered during the next task allocation.

[0064] In step S3, when determining the optimal resource range, first, the priorities of the logical resource pools are determined according to the magnitudes of the dynamic weights, and the priorities of the intelligent optical fiber distribution robots in each logical resource pool are determined according to the health status and the number of idle cores of the intelligent optical fiber distribution robots; then, a set number of logical resource pools and the intelligent optical fiber distribution robots they contain are selected in the order of priority as the optimal resource range, so that reinforcement learning performs core path selection based on the logical resource pools instead of in a huge and disordered set of physical cores, thereby greatly reducing the search space.

[0065] Among them, referring to the attached drawings of the specification Figure 3 , determining a scheduling path that meets the service requirements from the optimal resource range based on the search algorithm and the reinforcement learning algorithm includes the following steps: S301. Using a search algorithm based on network topology information, calculate a set of feasible paths from the source node to the destination node from the optimal resource range; wherein, if there is no path that meets the service requirements, expand the number of the selected logical resource pools and the intelligent optical fiber distribution robots they contain in the order of priority, and re-determine the optimal resource range until a path that meets the service requirements is found; S302. Obtain the state of the current optical fiber network, and select a path from the set of feasible paths as an action to execute using a greedy strategy; wherein, if the generated random number is less than the set exploration rate, randomly select a path from the set of feasible paths as an action to execute, and if the generated random number is greater than the set exploration rate, select the path with the maximum cumulative reward expectation value from the set of feasible paths as an action to execute; S303. Calculate the immediate reward after the action execution based on the set reward strategy, and calculate the cumulative reward expectation value after the action execution based on the immediate reward, the set learning rate and discount factor, the cumulative reward expectation value when the action is not executed, and the maximum cumulative reward expectation value among all feasible paths; S304. Perform iterative optimization of the path based on the calculated cumulative reward expectation value to determine a scheduling path that meets the service requirements.

[0066] In step S301, when using a search algorithm to screen a set of feasible paths that meet service requirements from the optimal resource range determined based on the priority to determine the path search path range, first, several logical resource pools with the highest weights are selected to ensure that the source node and the destination node belong to these logical resource pools. Then, for each selected logical resource pool, the intelligent optical fiber distribution robots participating in the path search are further determined according to the priority queue of the intelligent optical fiber distribution robots inside it, and the set of fiber cores managed by the selected intelligent optical fiber distribution robots is used as the initial range of the path search. Thus, from the huge physical fiber core set, a relatively small logical resource pool-intelligent optical fiber distribution robot-fiber core set that ensures fiber core quality and balanced business fiber core load is screened out based on the priorities of the logical resource pool and the intelligent optical fiber distribution robot, providing a reasonable starting space for subsequent path search. Furthermore, based on the network topology information, a search algorithm is used to calculate a set of feasible paths from the source node to the destination node from this fiber core set as the action space of reinforcement learning.

[0067] It should be noted that since a relatively small logical resource pool-intelligent optical fiber distribution robot-fiber core set is selected, there is a possibility that no feasible path can be found from this fiber core set. When this situation occurs, the logical resource pools are sequentially expanded downward from high to low according to the logical resource pool priority queue, and the optimal resource range is re-determined according to the intelligent optical fiber distribution robot queue priority until a feasible path is found or all logical resource pools are traversed.

[0068] In one embodiment, the A* algorithm is used to calculate the set of alternative paths. : When a service request is received and the start and end nodes from A to Z are determined, the possible fiber core paths from the start point to the end point are calculated based on the network topology information through the A* algorithm. Among them, the A* algorithm uses heuristic search to estimate the cost from node n to the target node on the basis of the breadth-first search algorithm. When selecting the next node to expand at each step, the A* algorithm will preferentially select the node with the smallest value.

[0069]

[0070] Among them, is the actual cost from the start node to node n, which is characterized by the normalized value of the sum of the physical laying lengths of all optical fiber lines from the start node to node n; is the estimated cost from node n to the target node, which is characterized by the normalized value of the shortest number of hops from node n to the target node; is the health loss cost of the intelligent optical fiber distribution robot corresponding to node n ( is the normalized score of the robot health, and the value range is 0-1); is the weight, and the sum of the three is 1. In this way, when calculating the link, both the influence of the optical cable length on the optical signal attenuation and the connection relationship of the subsequent nodes and the health of the currently selected intelligent optical fiber distribution robot are considered, and the influence of the intelligent optical fiber distribution robot's jump connection on the optical power attenuation is minimized as much as possible.

[0071] In steps S302 - S304, for each path in the set of feasible paths, its value is evaluated according to the current Q value, that is, learning records the long-term cumulative reward expectation of performing a certain action in a certain state by maintaining a Q value table. In the optical fiber distribution scenario, each row of the Q value table corresponds to the real-time state s at a certain moment t, and each column corresponds to a possible action (that is, the core connection path, calculated by the A* algorithm), and the evaluation value after this action is executed , and its update formula is:

[0072] Among them, is the Q value stored when the core connection path action is not executed in the state , is the Q value after executing the core connection path action , representing the long-term cumulative reward expectation of selecting this path in the current state ; represents the learning rate, and its value is between [0, 1]. It controls the degree of learning new information each time the Q value is updated. The closer it is to 1, the greater the influence of the newly obtained reward information on the update of the Q value; The closer it is to 0, the more the update of the Q value depends on previous experience; is the immediate reward obtained after executing the action , which is used to measure the direct effect of the action; is the discount factor, and its value range is between [0, 1]. It is used to measure the importance of future rewards. The closer it is to 1, the more it indicates that future rewards are valued, and the long-term influence of the current action on subsequent states and rewards will be considered; The closer it is to 0, the more it indicates that only immediate rewards are concerned; represents the new state migrated to after executing the action , reflecting the change of the state due to the execution of the action; represents all feasible core connection paths in the new state ; The largest Q value represents the maximum expected long-term cumulative reward that can be obtained by choosing the optimal path in the new state. Among them, in the initial stage, for each path in the set of feasible paths, the Q value is usually initialized to a default value, such as 0. This is because at this time, not enough experience has been accumulated and there is no clear judgment on the value of each path.

[0073] Among them, the state can be represented by a multi-dimensional vector, , is the network topology information, which is used to judge the physical connectivity of the network and includes the connection relationship between nodes, link length, and optical power attenuation. They are respectively composed of three N×N connection relationships (Connectivity), link lengths (LinkLength), and optical power attenuations (Attenuation) to form a three-dimensional matrix; represents the dynamic weights of each logical resource pool and is used to guide path search. It is a one-dimensional array. represents the predicted health values of each intelligent optical fiber distribution robot and is a one-dimensional array. represents the core occupancy status. The cores are numbered according to "resource pool - robot - core", and the core occupancy status is marked, which can quickly locate the available cores in the resource pool and is a two-dimensional array.

[0074] In the Q value calculation formula for each path, the action a represents the selection of a feasible optical fiber connection path from the source node to the destination node, that is, the selection of a specific core combination to establish an optical path connection. For example, selecting the core 1 of node A, passing through the core 5 of the intermediate node B, and finally connecting to the core 8 of the destination node C, this complete path selection is an action. In a certain state, when selecting an action (i.e., selecting a path) from the set of feasible paths, the following greedy strategy is adopted: with a probability of ε, randomly select a path from the set of feasible paths as the action; with a probability of 1 - ε, select the path with the highest Q value as the action:

[0075] Among them, is the action selected at time t; is the exploration rate, and its value range is between [0, 1]. In the initial stage of learning, is set to a relatively large value (such as 0.3), so that there are more opportunities to explore different actions; as learning progresses, gradually decreases and is more inclined to select the currently considered optimal action; It is a set of feasible paths calculated by the A* algorithm according to the network topology and service requests (start and end nodes, optical power requirements, etc.) that meet the optical power index requirements. In a complex optical fiber network, the A* algorithm uses a heuristic function to estimate the cost from the current node to the target node, thus quickly screening out feasible paths; The Q value for executing action a in state represents the long-term cumulative reward expectation for selecting action a in state

[0076] Assume the current state is , and path is selected from the set of feasible paths. After executing this path selection action, according to information such as whether the connection is successfully established and whether the optical power meets the requirements, the immediate reward is obtained. Then, according to the Q-value update formula, the new Q value is calculated. For example, if the connection is successfully established and the optical power meets the requirements, is positive (such as +1); if the connection fails or the optical power does not meet the standard, is negative (such as -1). is the learning rate, which controls the influence degree of the new reward information on the Q-value update; is the discount factor, which measures the importance of future rewards. represents the maximum Q value among all possible actions (i.e., all feasible fiber core connection paths) in the new state . Through continuous iterative updates, the Q value gradually reflects the value of each path in different states. With the accumulation of experience, the master control node evaluates the path value according to the updated Q value and preferentially selects the path with a high Q value to achieve a better decision-making effect.

[0077] Among them, the immediate reward after the execution of the action is calculated based on the set reward strategy, and the reward strategy can be set according to one or more indicators among connection quality, fiber core health, service target, and path efficiency. For example, according to the selected action, a fiber optic patch cord instruction is sent to the specified intelligent fiber optic distribution robot, and the intelligent fiber optic distribution robot performs the fiber core connection operation. After the connection is completed, if a connection that meets the optical power index requirements is successfully established, an immediate reward of +1 is given; if the connection fails, an immediate reward of -1 is given; if the connection is successfully established but exceeds the optical power index threshold, a small negative reward (such as -0.5) is given. It is specifically set according to the actual application, and this application does not limit and fix it.

[0078] ​The above-mentioned state information acquisition, action selection, action execution, reward calculation, and policy update are repeated iteratively. As the number of iterations increases, the policy is continuously learned and optimized, and a better core path selection method is gradually found. Thus, by combining the A* algorithm with Q-value reinforcement learning, network topology information can be more effectively utilized in the action selection stage to find a better core path for service requests, breaking through the limitations of traditional path selection methods and no longer blindly searching or relying solely on single-factor decision-making. At the same time, the learning ability of reinforcement learning is used to continuously optimize the policy, improve the SLA (Service-Level Agreement) satisfaction rate and resource utilization rate to adapt to the dynamic changes of the network state.

[0079] Then, the physical core is abstracted into a logical resource pool, and dynamic weight partitioning scheduling is adopted to achieve optimal resource allocation. It breaks through the physical limitations of traditional resource allocation and is no longer restricted by the core range managed by a single robot. By calculating the dynamic weight for each logical resource pool based on the real-time load and core health, service requests can be accurately allocated to the optimal robot management area, achieving robot load balancing and overload avoidance. Compared with the traditional fixed resource allocation mode, the utilization efficiency of optical fiber resources is greatly improved, and the problem of rigid resource allocation is effectively solved in the high-density computer room scenario, improving the overall operation and maintenance efficiency. In addition, reinforcement learning is combined with the A* algorithm to generate a scheduling path that meets the requirements of optical power and service level. Using this method, the limitations of traditional path selection methods are broken through, and no longer blindly search or rely solely on single-factor decision-making. And with the accumulation of experience, the path selection policy is continuously optimized, the SLA satisfaction rate and resource utilization rate are improved, and it can better adapt to the dynamic changes of the network state to achieve efficient service scheduling in a complex optical fiber network environment.

[0080] Furthermore, when it is detected that the network, device, or link state changes or the service request volume fluctuates beyond the set threshold, the dynamic weight is recalculated and the scheduling path is determined again.

[0081] (1) When the network state changes, such as a sudden failure in a robot management area, immediately recalculate the dynamic weights of all logical resource pools. During the recalculation process, mark the cores in the failure area as unavailable and reduce the weight of the logical resource pool where they are located. At the same time, re-evaluate the set of feasible paths for all service requests. For service requests originally allocated to the logical resource pool related to the failure area, based on the updated network topology and weight information, use the A* algorithm again to calculate the feasible paths. During the calculation process, avoid selecting the cores in the failure area and update the Q-value table according to the new set of feasible paths.

[0082] (2) When the service request volume fluctuates significantly, dynamically adjust the weight coefficients of the load and health according to the length of the real-time service request queue and the load conditions of each logical resource pool , . If the request volume is large, increase the proportion of the load factor in the weight calculation, and preferentially select the logical resource pool with more idle cores; if the request volume is small, appropriately increase the weight of the health factor, and preferentially select the area with high core health. Then, according to the updated dynamic weights, adjust the action selection strategy. When selecting a path, pay more attention to the logical resource pools with large weight changes, and preferentially search for feasible paths in these areas to improve the resource allocation efficiency.

[0083] Among them, when multiple service requests are received, the requests are sorted according to the priority to form a priority queue. When multiple high-priority service requests compete for the same resource, first check whether there are other available alternative resources (such as spare cores, idle cores in other logical resource pools). If there are alternative resources, allocate the service requests for the competing resources to the alternative resources. If there are no alternative resources, then according to the waiting time and resource requirements of the services, adopt the following strategy: for services with a long waiting time and relatively small resource requirements, allocate resources preferentially; for services with similar waiting times, further subdivide according to the service priority, and preferentially meet the requirements of higher-priority services.

[0084] (3) When a certain robot fails, stop allocating new service requests to this robot, and mark the cores it manages as unavailable. Recalculate the dynamic weights of all logical resource pools, and reduce the weight of the logical resource pool where this robot is located. For the fiber hopping services being executed on this robot, if other robots have spare idle cores, migrate the services to the corresponding robots, re-plan the fiber connection paths, and update the service connection relationships. If other robots do not have spare idle cores, reschedule the services according to the service priority and importance. For high-priority services, try to find alternative paths in the management areas of other robots; for low-priority services, they can wait temporarily or be processed according to the allowed interruption time of the services. At the same time, notify the maintenance personnel to repair the faulty robot. After the robot is repaired, include it in the resource management system again and resume normal service scheduling.

[0085] (4) When a core failure is detected, immediately search for spare cores according to the information of the robot management area and logical resource pool where the faulty core is located. If there are spare cores in this robot management area, send an instruction to the corresponding robot to switch the service of the faulty core to the spare core. At the same time, update the core occupancy and service connection relationships to ensure the continuity of the service. If there are no spare cores in this area, the main control node recalculates the dynamic weights, searches for available cores in other logical resource pools, uses the A* algorithm to re-plan the service path to avoid the area where the faulty core is located, and sends new fiber hopping instructions to the relevant robots. After the service switch is completed, mark the faulty core and arrange for the maintenance personnel to carry out maintenance.

[0086] It can be seen that an intelligent optical fiber distribution collaborative scheduling method provided by the application integrates federated learning, dynamic weight scheduling, and reinforcement learning, constructs a three-layer collaborative architecture, can automatically perform resource allocation and path selection decisions according to different service requirements, reduce manual intervention, reduce human errors, and achieve intelligent operation and maintenance management. At the same time, the dynamic weight partitioning mechanism and the reinforcement learning algorithm can flexibly adjust the strategy according to the network state changes to adapt to the dynamic changes of the network.

[0087] Based on the same inventive concept, an intelligent optical fiber distribution collaborative scheduling device is also provided in an embodiment of the present application. Since the principle of solving problems by the device in the embodiment of the present application is similar to that of the above-mentioned intelligent optical fiber distribution collaborative scheduling method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0088] As shown in the attached Figure 4 description, an embodiment of the present application also provides an intelligent optical fiber distribution collaborative scheduling device, which is applied to a cluster composed of multiple intelligent optical fiber distribution robots. The device includes: A federated learning health prediction module 401, configured to train a global LSTM model for predicting the health status of intelligent optical fiber distribution robots based on federated learning to obtain global model parameters, and load the global model parameters into the local LSTM models of each intelligent optical fiber distribution robot to evaluate the health of each intelligent optical fiber distribution robot; A dynamic weight partitioning module 402, configured to abstract the fiber cores of different intelligent optical fiber distribution robots into a logical resource pool, and calculate its dynamic weight according to the load condition of each logical resource pool and the health of the intelligent optical fiber distribution robots included; A reinforcement learning path matching module 403, configured to determine an optimal resource range based on the dynamic weight, and determine a scheduling path that meets service requirements from the optimal resource range based on a search algorithm and a reinforcement learning algorithm.

[0089] In an embodiment, the device further includes: An adaptive operation and maintenance module, configured to recalculate the dynamic weight and determine the scheduling path when it is detected that the network, device, and link states change or the service request volume fluctuates beyond a set threshold.

[0090] In one embodiment, the federated learning health prediction module 401 trains a global LSTM model for predicting the health status of the intelligent optical fiber distribution robot based on federated learning to obtain global model parameters, including: initializing the global LSTM model and sending the initialized global model parameters to the local LSTM models of each intelligent optical fiber distribution robot; preprocessing the device data collected by the intelligent optical fiber distribution robot to construct time series data of different modalities, and training the local LSTM model based on the time series data; performing parameter alignment compensation and parameter aggregation strategies on the global LSTM model using the local training parameters obtained by each intelligent optical fiber distribution robot to obtain updated global model parameters.

[0091] In one embodiment, the device data includes optical power loss, the number of pluggings and unpluggings of the optical fiber connector, the running mileage of the lead screw, and environmental data. The federated learning health prediction module 401 preprocesses the device data collected by the intelligent optical fiber distribution robot to construct time series data of different modalities, including: statistically counting the number of pluggings and unpluggings within each window based on the sliding window mechanism, and dividing it into a low-frequency plugging and unplugging scenario and a high-frequency plugging and unplugging scenario according to a preset threshold value; wherein, in the low-frequency plugging and unplugging scenario, extreme value normalization is performed on the optical power loss; in the high-frequency plugging and unplugging scenario, standard deviation normalization is performed on the optical power loss; calculating the cumulative fatigue degree based on the number of pluggings and unpluggings, the running mileage of the lead screw, and the time decay effect, which is used to measure the cumulative pressure and loss endured by the device during use; obtaining the time stamp of each plugging and unplugging operation, and converting it into a first time code and a second time code according to the periodicity of the plugging and unplugging of the intelligent optical fiber distribution robot; aligning the normalized optical power loss, cumulative fatigue degree, number of pluggings and unpluggings, first time code, second time code, and environmental data according to time steps to form a two-dimensional input tensor as time series data of different modalities.

[0092] In one embodiment, the reinforcement learning path matching module 403 determines the optimal resource range based on the dynamic weight, including: determining the priority of the logical resource pool according to the magnitude sorting of the dynamic weight, and determining the priority of each intelligent optical fiber distribution robot in the logical resource pool according to the health degree and the number of idle fiber cores of the intelligent optical fiber distribution robot; selecting a set number of logical resource pools and the intelligent optical fiber distribution robots they contain in the order of priority as the optimal resource range.

[0093] In one embodiment, the reinforcement learning path matching module 403 determines a scheduling path that meets service requirements from the optimal resource range based on a search algorithm and a reinforcement learning algorithm, including: calculating a set of feasible paths from the source node to the destination node from the optimal resource range using a search algorithm based on network topology information; wherein, if there is no path that meets service requirements, expand the selected logical resource pool and the number of intelligent fiber optic distribution robots it contains in the order of priority, and re-determine the optimal resource range until a path that meets service requirements is found; obtain the state of the current fiber optic network, and select a path from the set of feasible paths as an action to execute using a greedy strategy; wherein, if the generated random number is less than the set exploration rate, randomly select a path from the set of feasible paths as an action to execute, if the generated random number is greater than the set exploration rate, select the path with the maximum cumulative reward expectation value from the set of feasible paths as an action to execute; calculate the immediate reward after the action execution based on the set reward strategy, and calculate the cumulative reward expectation value after the action execution based on the immediate reward, the set learning rate and discount factor, the cumulative reward expectation value when the action is not executed, and the maximum cumulative reward expectation value among all feasible paths; perform iterative optimization of the path based on the calculated cumulative reward expectation value to determine a scheduling path that meets service requirements. The search algorithm uses the A* algorithm, and the state of the fiber optic network includes network topology information, the dynamic weight of the logical resource pool, and the health of the intelligent fiber optic distribution robot; the network topology information includes the connection relationship between nodes, the link length, and the optical power attenuation.

[0094] The intelligent fiber optic distribution collaborative scheduling device described in this application is applied to a cluster composed of multiple intelligent fiber optic distribution robots. The global LSTM model for predicting the health status of intelligent fiber optic distribution robots is trained through a federated learning health prediction module based on federated learning to obtain global model parameters, and the global model parameters are loaded into the local LSTM models of each intelligent fiber optic distribution robot to evaluate the health of each intelligent fiber optic distribution robot; the cores of different intelligent fiber optic distribution robots are abstracted into logical resource pools through a dynamic weight partitioning module, and the dynamic weight of each logical resource pool is calculated according to the load situation of each logical resource pool and the health of the intelligent fiber optic distribution robots it contains; the reinforcement learning path matching module determines the optimal resource range based on the dynamic weight, and determines a scheduling path that meets service requirements from the optimal resource range based on a search algorithm and a reinforcement learning algorithm. Thus, by integrating federated learning, dynamic weight scheduling, and reinforcement learning, a three-layer collaborative architecture is constructed, which can automatically perform resource allocation and path selection decisions according to different service requirements, reduce manual intervention, reduce human errors, and achieve intelligent operation and maintenance management.

[0095] Based on the same concept of the present invention, as shown in the accompanying Figure 5As shown, the structure of an electronic device 500 provided by an embodiment of the present application. The electronic device 500 includes: at least one processor 501, at least one network interface 504 or other user interfaces 503, a memory 505, and at least one communication bus 502. The communication bus 502 is used to realize the connection and communication between these components. The electronic device 500 optionally includes a user interface 503, including a display (such as a touch screen, LCD, CRT, holographic imaging, or projector, etc.), a keyboard, or a pointing device (such as a mouse, trackball, touchpad, or touch screen, etc.).

[0096] The memory 505 may include a read-only memory and a random access memory, and provide instructions and data to the processor 501. A part of the memory 505 may also include a non-volatile random access memory (NVRAM).

[0097] In some embodiments, the memory 505 stores the following elements, executable modules, or data structures, or subsets thereof, or extended sets thereof: An operating system 5051, including various system programs, used to implement various basic services and process hardware-based tasks; An application program module 5052, including various application programs, such as a launcher, a media player, a browser, etc., used to implement various application services.

[0098] In the embodiment of the present application, by invoking the programs or instructions stored in the memory 505, the processor 501 is used to execute the steps of a method for intelligent fiber distribution collaborative scheduling.

[0099] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps in a method for intelligent fiber distribution collaborative scheduling.

[0100] Specifically, the storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the storage medium is run, it can realize the dynamic scheduling of fiber resources within the cluster through a three-layer collaborative architecture of federated learning prediction - dynamic weight scheduling - reinforcement learning decision-making.

[0101] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some communication interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0102] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0103] In addition, each functional unit in the embodiments provided in this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0104] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0105] Finally, it should be noted that the above embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. All should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. An intelligent optical fiber distribution collaborative scheduling method, characterized in that, Applied to a cluster composed of multiple intelligent optical fiber distribution robots, the method includes the following steps: Based on federated learning, train a global LSTM model for predicting the health status of intelligent optical fiber distribution robots to obtain global model parameters, and load the global model parameters into the local LSTM models of each intelligent optical fiber distribution robot to evaluate the health of each intelligent optical fiber distribution robot; Abstract the fiber cores of different intelligent optical fiber distribution robots into a logical resource pool, and calculate its dynamic weight according to the load condition of each logical resource pool and the health of the intelligent optical fiber distribution robots included; Determine the optimal resource range based on the dynamic weight, and determine the scheduling path that meets the service requirements from the optimal resource range based on the search algorithm and the reinforcement learning algorithm.

2. The intelligent optical fiber distribution collaborative scheduling method according to claim 1, wherein The method further includes the following steps: When it is monitored that the network, device, and link status change or the service request volume fluctuates beyond the set threshold, recalculate the dynamic weight and determine the scheduling path.

3. The intelligent optical fiber distribution collaborative scheduling method according to claim 1, wherein The training of the global LSTM model for predicting the health status of intelligent optical fiber distribution robots based on federated learning to obtain global model parameters includes the following steps: Initialize the global LSTM model, and send the initialized global model parameters to the local LSTM models of each intelligent optical fiber distribution robot; Preprocess the device data collected by the intelligent optical fiber distribution robot, construct time series data of different modalities, and train the local LSTM model based on the time series data; Execute the parameter alignment compensation and parameter aggregation strategy on the global LSTM model using the local training parameters obtained by each intelligent optical fiber distribution robot to obtain the updated global model parameters.

4. The intelligent optical fiber distribution collaborative scheduling method according to claim 3, wherein, Wherein, The device data includes optical power loss, the number of pluggings and unpluggings of the optical fiber connector, the running mileage of the lead screw, and environmental data. The preprocessing of the device data collected by the intelligent optical fiber distribution robot to construct time series data of different modalities includes the following steps: Based on the sliding window mechanism, count the number of pluggings and unpluggings within each window, and divide it into a low-frequency plugging and unplugging scenario and a high-frequency plugging and unplugging scenario according to a preset threshold value; wherein, in the low-frequency plugging and unplugging scenario, perform extreme value normalization processing on the optical power loss; in the high-frequency plugging and unplugging scenario, perform standard deviation normalization processing on the optical power loss; Calculate the cumulative fatigue degree based on the number of pluggings and unpluggings, the running mileage of the lead screw, and the time decay effect, which is used to measure the cumulative pressure and loss endured by the device during use; Obtain the timestamp of each plugging and unplugging operation, and convert it into a first time code and a second time code according to the periodicity of the plugging and unplugging of the intelligent optical fiber distribution robot; Align the normalized optical power loss, cumulative fatigue degree, number of pluggings and unpluggings, first time code, second time code, and environmental data according to time steps to form a two-dimensional input tensor as time series data of different modalities.

5. The intelligent optical fiber distribution collaborative scheduling method according to claim 1, wherein The determination of the optimal resource range based on the dynamic weight includes the following steps: Determine the priority of the logical resource pool according to the sorting of the magnitudes of the dynamic weights, and determine the priorities of the individual intelligent optical fiber distribution robots in the logical resource pool according to the health and the number of idle cores of the intelligent optical fiber distribution robot; Select a set number of logical resource pools and the intelligent optical fiber distribution robots they contain in the order of priority as the optimal resource range.

6. The intelligent optical fiber distribution collaborative scheduling method according to claim 5, wherein, The determining of the scheduling path that meets the service requirements from the optimal resource range based on the search algorithm and the reinforcement learning algorithm includes the following steps: Calculate a set of feasible paths from the source node to the destination node from the optimal resource range using the search algorithm based on the network topology information; wherein, if there is no path that meets the service requirements, expand the number of the selected logical resource pools and the intelligent optical fiber distribution robots they contain in the order of priority, and re-determine the optimal resource range until a path that meets the service requirements is found; Obtain the state of the current optical fiber network, and select a path from the set of feasible paths as the action to be executed using the greedy strategy; wherein, if the generated random number is less than the set exploration rate, randomly select a path from the set of feasible paths as the action to be executed, and if the generated random number is greater than the set exploration rate, select the path with the maximum cumulative reward expectation value from the set of feasible paths as the action to be executed; Calculate the immediate reward after the action is executed based on the set reward strategy, and calculate the cumulative reward expectation value after the action is executed based on the immediate reward, the set learning rate and discount factor, the cumulative reward expectation value when the action is not executed, and the maximum cumulative reward expectation value among all feasible paths; Perform iterative optimization of the path based on the calculated cumulative reward expectation value to determine the scheduling path that meets the service requirements.

7. The intelligent optical fiber distribution collaborative scheduling method according to claim 6, wherein, Wherein, The search algorithm uses the A* algorithm, and the state of the optical fiber network includes network topology information, the dynamic weights of the logical resource pools, and the health of the intelligent optical fiber distribution robots; the network topology information includes the connection relationship between nodes, the link length, and the optical power attenuation.

8. An intelligent optical fiber distribution collaborative scheduling device, characterized in that, Applied to a cluster composed of multiple intelligent optical fiber distribution robots, the device includes: A federated learning health prediction module, which is used to train a global LSTM model for predicting the health status of intelligent optical fiber distribution robots based on federated learning to obtain global model parameters, and load the global model parameters into the local LSTM models of the individual intelligent optical fiber distribution robots to evaluate the health of the individual intelligent optical fiber distribution robots; A dynamic weight partitioning module, which is used to abstract the cores of different intelligent optical fiber distribution robots into logical resource pools, and calculate their dynamic weights according to the load conditions of each logical resource pool and the health of the intelligent optical fiber distribution robots it contains; A reinforcement learning path matching module, which is used to determine the optimal resource range based on the dynamic weights, and determine the scheduling path that meets the service requirements from the optimal resource range based on the search algorithm and the reinforcement learning algorithm.

9. An electronic device, characterized in that, Include: A processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of an intelligent optical fiber wiring collaborative scheduling method according to any one of claims 1 to 7 are performed.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by the processor, the steps of an intelligent optical fiber wiring collaborative scheduling method according to any one of claims 1 to 7 are performed.

Citation Information

Patent Citations

  • Federal learning method and device, equipment and storage medium

    CN113516250A

  • Task execution method and device of optical fiber wiring robot, electronic equipment and medium

    CN116582180A

  • Adaptive asynchronous federated learning method and system based on deep reinforcement learning

    CN118586474A

  • Wireless communication and computing resource collaborative optimization method and system

    CN119512766A

  • Unstructured database federated learning collaboration method and system based on swarm intelligence

    CN119669433A