Intelligent optical fiber distribution scheduling method and device based on reinforcement learning decision

By abstracting the fiber core of the optical fiber network into a logical resource pool and using a reinforcement learning algorithm to dynamically generate the optimal fiber core path, the limitations of traditional optical fiber network resource scheduling are overcome, and efficient fiber core path matching and service response are achieved.

CN120602358AActive Publication Date: 2025-09-05CHINA TRANSPORT INFORMATION TECH GRP CO LTD +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510748481.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-05
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

In application scenarios where high service provisioning and recovery speeds are required, traditional fiber optic network resource management lacks a global resource pooling mechanism, resulting in fiber core resource scheduling requiring manual intervention or relying on static routing strategies, making it impossible to respond to critical business needs in a timely manner.

Method used

The fiber cores in the optical fiber network are abstracted into a logical resource pool, and the priority and dynamic weight are determined through the reinforcement learning algorithm. The search algorithm and the A-star algorithm are combined to dynamically generate the optimal fiber core path, matching business needs with resource status in real time.

Benefits of technology

It realizes dynamic resource configuration of optical fiber network, improves business response efficiency, reduces business waiting time, and improves resource utilization and SLA compliance rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602358A_ABST
    Figure CN120602358A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent optical fiber distribution scheduling method and device based on reinforcement learning decision, and belongs to the technical field of optical fiber communication, and the method comprises the steps: abstracting fiber cores of different intelligent optical fiber distribution robots in a cluster into a logic resource pool; the priority of the logic resource pool and the priority of the intelligent optical fiber distribution robot contained in the logic resource pool are determined through dynamic weight partition, the physical limitation of traditional resource distribution is broken through, and the physical limitation is not limited by the fiber core range managed by a single intelligent optical fiber distribution robot any more; through the reinforcement learning technology, the service demand and the resource state are matched in real time, the optimal fiber core path is dynamically generated, the fiber core path meeting the service requirement can be quickly found in an application scene with high requirements for service opening and recovery speed, the service response efficiency is improved, and the service waiting time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of optical fiber communication technology, and more specifically, to an intelligent optical fiber distribution scheduling method and device based on reinforcement learning decision-making. Background Art

[0002] With the surge in demand for services such as 5G and data centers, the scale of optical cable networks continues to expand, with the number of network nodes and cable mileage both growing exponentially, creating an urgent need for intelligent operations and maintenance (O&M). Currently, optical fiber network O&M is still primarily manual, with patching configuration, troubleshooting, and other tasks relying on on-site manual labor, resulting in low efficiency and susceptible to human error. As the scale of optical cable networks continues to expand, intelligent fiber optic patching robots are gradually replacing traditional manual operations. Using robotic arms, they perform automated patching and fiber core quality monitoring, enabling dynamic resource allocation and intelligent management of optical fiber networks. These robots can remotely perform fiber core insertion and removal, path switching, and performance testing, significantly improving O&M efficiency.

[0003] However, in application scenarios with high requirements for service provisioning and recovery speed, fiber core resources in traditional solutions are independently managed by single-machine robots, lacking a global resource pooling mechanism. Cross-robot scheduling requires manual intervention or fixed routing strategies, that is, relying on preset static routing tables or single-machine local searches, resulting in the problem of critical services being unable to be scheduled in a timely manner. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide an intelligent optical fiber distribution scheduling method and device based on reinforcement learning decision-making, which can match business needs and resource status in real time through reinforcement learning technology, dynamically generate the optimal fiber core path, and improve business response efficiency.

[0005] An embodiment of the present application provides an intelligent fiber optic wiring scheduling method based on reinforcement learning decision-making, which is applied to a cluster composed of multiple intelligent fiber optic wiring robots. The method includes the following steps:

[0006] Abstract the fiber cores connected to different intelligent fiber distribution robots in the cluster into a logical resource pool;

[0007] Determining the priority of the logical resource pool and the intelligent fiber optic wiring robot contained therein;

[0008] A scheduling path that meets business requirements is determined based on a reinforcement learning algorithm. Specifically, a search algorithm is first used to screen a set of feasible paths that meet business requirements in a logical resource pool determined based on the priority. Then, the current state of the optical fiber network is obtained, and a path is selected from the set of feasible paths as an action to be executed. An immediate reward and an expected value of a cumulative reward are calculated based on the execution result. Finally, based on the calculated expected value of the cumulative reward, the path is iteratively optimized to determine a scheduling path that meets business requirements.

[0009] In some embodiments, the step of abstracting the fiber cores connected to different intelligent fiber distribution robots in the cluster into a logical resource pool includes the following steps:

[0010] Configure logical identification for the fiber core range managed by each intelligent fiber optic wiring robot;

[0011] The dynamic weight of each logical resource pool is calculated based on its load and average fiber core health.

[0012] In some embodiments, determining the priority of the logical resource pool and the intelligent fiber optic distribution robot contained therein comprises the following steps:

[0013] Determining the priority of the logical resource pool according to the size of the dynamic weight;

[0014] The priority of each intelligent fiber optic patching robot in the logical resource pool is determined based on the health status and number of idle fiber cores of the intelligent fiber optic patching robot.

[0015] In some embodiments, the step of using a search algorithm to screen a set of feasible paths that meet business requirements in the logical resource pool determined based on the priority includes the following steps:

[0016] Selecting a set number of logical resource pools and the intelligent fiber optic wiring robots contained therein in order of priority to obtain a managed fiber core set;

[0017] Calculating a set of feasible paths from a source node to a destination node from the set of fiber cores using an A-star algorithm based on network topology information;

[0018] Determine whether there is a path in the feasible path set that meets the business requirements. If so, use the feasible path set as the action space for reinforcement learning; if not, expand the selected logical resource pool and the number of intelligent fiber optic distribution robots it contains in order of priority, and recalculate the feasible path set until a path that meets the business requirements is found.

[0019] In some embodiments, the state of the optical fiber network includes network topology information, dynamic weights of logical resource pools, and health of intelligent optical fiber distribution robots; the network topology information includes connection relationships between nodes, link lengths, and optical power attenuation.

[0020] In some embodiments, selecting a path from the set of feasible paths as an action to be executed, and calculating the immediate reward and the expected value of the cumulative reward based on the execution result, includes the following steps:

[0021] A greedy strategy is used to select a path from the set of feasible paths as an action to be executed; if the generated random number is less than the set exploration rate, a path is randomly selected from the set of feasible paths as an action to be executed; if the generated random number is greater than the set exploration rate, the path with the largest expected cumulative reward is selected from the set of feasible paths as an action to be executed;

[0022] The immediate reward after the action is executed is calculated based on the set reward strategy, and the expected cumulative reward after the action is executed is calculated based on the immediate reward, the set learning rate and discount factor, the expected cumulative reward when the action is not executed, and the maximum expected cumulative reward among all feasible paths.

[0023] In some embodiments, the reward strategy is set according to one or more indicators of connection quality, fiber core health, service goals, and path efficiency.

[0024] In some embodiments, an intelligent fiber optic wiring scheduling device based on reinforcement learning decision-making is further provided, which is applied to a cluster composed of multiple intelligent fiber optic wiring robots, and the device includes:

[0025] A building block for abstracting the fiber cores connected to different intelligent fiber distribution robots in the cluster into a logical resource pool;

[0026] A determination module, configured to determine the priority of the logical resource pool and the intelligent optical fiber distribution robot contained therein;

[0027] A path matching module is used to determine a scheduling path that meets business requirements based on a reinforcement learning algorithm. A search algorithm is first used to screen a set of feasible paths that meet business requirements from a logical resource pool determined based on the priority level. The current state of the optical fiber network is then obtained, and a path is selected from the set of feasible paths as an action to be executed. An immediate reward and an expected cumulative reward are calculated based on the execution result. Finally, based on the calculated expected cumulative reward, the path is iteratively optimized to determine a scheduling path that meets business requirements.

[0028] In some embodiments, an electronic device is also provided, comprising: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the steps of any one of the above-mentioned intelligent fiber optic distribution scheduling methods based on reinforcement learning decision-making are performed.

[0029] In some embodiments, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned intelligent fiber optic distribution scheduling methods based on reinforcement learning decision-making are executed.

[0030] The present application describes an intelligent fiber optic patching scheduling method and device based on reinforcement learning decision-making, which is applied to a cluster composed of multiple intelligent fiber optic patching robots. The method abstracts the fiber cores of different intelligent fiber optic patching robots connected to the cluster into a logical resource pool; determines the priority of the logical resource pool and the intelligent fiber optic patching robots contained therein; and determines a scheduling path that meets business requirements based on a reinforcement learning algorithm. A search algorithm is first used to screen a set of feasible paths that meet business requirements in the logical resource pool determined based on the priority; then, the current state of the optical fiber network is obtained, and a path is selected from the set of feasible paths as an action to be executed. Based on the execution result, an immediate reward and an expected cumulative reward value are calculated; finally, based on the calculated expected cumulative reward value, the path is iteratively optimized to determine a scheduling path that meets business requirements. Thus, by combining reinforcement learning with dynamic weighted scheduling, business needs and resource status are matched in real time, and the optimal fiber core path is dynamically generated. For application scenarios with high service provisioning and recovery speed requirements, the method can quickly find a fiber core path that meets business requirements, improve business response efficiency, and reduce business waiting time. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0032] Figure 1 A flowchart of the intelligent optical fiber distribution scheduling method based on reinforcement learning decision-making according to an embodiment of the present application is shown;

[0033] Figure 2 A flowchart of determining the priority of the logical resource pool and the intelligent fiber optic wiring robot contained therein according to an embodiment of the present application is shown;

[0034] Figure 3 A flowchart of selecting a path from the set of feasible paths as an action to be executed and calculating the instant reward and the expected value of the cumulative reward based on the execution result is shown in an embodiment of the present application;

[0035] Figure 4 A schematic diagram of the structure of an intelligent optical fiber distribution and scheduling device based on reinforcement learning decision-making according to an embodiment of the present application is shown;

[0036] Figure 5 A schematic structural diagram of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.

[0038] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.

[0039] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.

[0040] In view of the technical problems raised by the background technology, the present application provides an intelligent optical fiber distribution scheduling method and device based on reinforcement learning decision-making, which can match business needs and resource status in real time through reinforcement learning technology, dynamically generate the optimal fiber core path, and improve business response efficiency.

[0041] See the instructions attached Figure 1 The present application provides an intelligent fiber optic wiring scheduling method based on reinforcement learning decision-making, which is applied to a cluster composed of multiple intelligent fiber optic wiring robots, including the following steps:

[0042] S1. Abstract the fiber cores connected to different intelligent fiber distribution robots in the cluster into a logical resource pool;

[0043] S2. Determine the priority of the logical resource pool and the intelligent fiber optic wiring robot contained therein;

[0044] S3. Determine a scheduling path that meets business requirements based on a reinforcement learning algorithm; wherein, a search algorithm is first used to screen a set of feasible paths that meet business requirements in a logical resource pool determined based on the priority; then, the current state of the optical fiber network is obtained, and a path is selected from the set of feasible paths as an action to be executed, and an immediate reward and an expected value of the cumulative reward are calculated based on the execution result; finally, the path is iteratively optimized based on the calculated expected value of the cumulative reward to determine a scheduling path that meets business requirements.

[0045] In step S1, the main task is to construct a logical resource pool. Specifically, the optical fiber cores connected to different intelligent fiber optic patching robots are abstracted into logical resource pools. A unique logical identifier is assigned to the range of fiber cores managed by each intelligent fiber optic patching robot. For example, logical resource pool 1 includes two intelligent fiber optic patching robots, A and B. Intelligent fiber optic patching robot A manages fiber cores 1-288; intelligent fiber optic patching robot B manages fiber cores 289-480. The logical resource pool identifiers corresponding to intelligent fiber optic patching robots A and B are Pool_1; logical resource pool 2 includes one intelligent fiber optic patching robot, C. Intelligent fiber optic patching robot C manages fiber cores 481-768, and the corresponding logical resource pool identifier is Pool_2. The logical identifiers of the fiber cores managed by all intelligent fiber optic patching robots are then integrated to form a global logical resource pool.

[0046] When a business request requires the selection of two intelligent fiber optic wiring robots to connect to different fiber cores, the dynamic weight W is calculated based on the real-time load of each logical resource pool and the average health of the fiber core. i , to balance the load of the intelligent fiber optic wiring robot and prioritize the assignment of tasks to the fiber core area managed by the intelligent fiber optic wiring robot with high health.

[0047]

[0048] Among them, W i is the dynamic weight of the i-th logical resource pool. The higher the weight, the more priority the logical resource pool will be selected when allocating tasks. ω1 and ω2 are weight coefficients used to adjust the relative importance of load and average core health, ω1+ω2=1. IdleCores i TotalCore is the number of currently idle cores in the i-th logical resource pool, reflecting the load of the logical resource pool. The more idle cores, the lower the load. i is the total number of cores in the i-th logical resource pool; Hpool i The average fiber core health of the i-th logical resource pool is obtained by averaging the health scores of all fiber cores in the logical resource pool. It reflects the health status of the fiber cores. The higher the score, the better the fiber core health.

[0049]

[0050] Where n is the number of intelligent fiber optic wiring robots in the logical resource pool; RH i Represents the health score of the i-th intelligent fiber optic wiring robot with a value range of 0-1, RH i =Hscore i / 100;IC i The number of idle fiber cores connected to the i-th intelligent fiber optic patching robot. This parameter reflects the remaining resources of the intelligent fiber optic patching robot that can be used to assign tasks. The more idle fiber cores there are, the more likely this intelligent fiber optic patching robot will be considered for the next task assignment.

[0051] Among them, Hscore i is the predicted health value of the i-th intelligent fiber optic wiring robot, expressed in percentage. In practical applications, there are many models to evaluate the health of various types of intelligent fiber optic robot equipment and operating conditions. For example, a trained LSTM model is used to analyze the currently collected operating data of the intelligent fiber optic wiring robot. This is not a key technical point of this patent and will not be elaborated here.

[0052] In step S2, see the attached Figure 2 , the determining of the priority of the logical resource pool and the intelligent fiber optic wiring robot contained therein comprises the following steps:

[0053] S201, determining the priority of the logical resource pool according to the size of the dynamic weight;

[0054] S202: Determine the priority of each intelligent fiber optic distribution robot in the logical resource pool according to the health status and the number of idle fiber cores of the intelligent fiber optic distribution robot.

[0055] Specifically, in step S201 and step S202, according to the dynamic weight W i The sorting results are used to generate a logical resource pool priority queue. If the logical resource pool contains multiple intelligent fiber optic wiring robots, the priority queue is further sorted according to the health level RH of the intelligent fiber optic wiring robots. i and the number of idle fiber cores IC i, generating a priority queue for the intelligent fiber optic patching robots within this logical resource pool. During subsequent route calculations, the logical resource pool with the highest weight is prioritized for connection with the fiber cores managed by the intelligent fiber optic patching robots. This step S2 identifies one or more high-priority logical resource pools and the relatively superior intelligent fiber optic patching robots within these logical resource pools. This defines the resource selection range for subsequent reinforcement learning path matching, allowing reinforcement learning to perform fiber core path selection based on the logical resource pools rather than a large and disordered set of physical fiber cores, significantly reducing the search space.

[0056] In step S3, a search algorithm is used to screen the logical resource pools determined based on the priorities for feasible paths that meet service requirements. To determine the path search scope, several logical resource pools with the highest weights are first selected to ensure that the source and destination nodes belong to these logical resource pools. Then, for each selected logical resource pool, the intelligent fiber optic patching robot participating in the path search is determined based on its internal intelligent fiber optic patching robot priority queue. The set of fiber cores managed by the selected intelligent fiber optic patching robot serves as the initial scope for the path search. Thus, based on the logical resource pools and the priorities of the intelligent fiber optic patching robots, a relatively small set of logical resource pools, intelligent fiber optic patching robots, and fiber cores is selected from the vast set of physical fiber cores, ensuring fiber core quality and balanced service core load. This provides a reasonable starting space for subsequent path search. Based on network topology information, a search algorithm is then used to calculate a set of feasible paths from the source node to the destination node from this set of fiber cores, which serves as the action space for reinforcement learning.

[0057] It should be noted that because a smaller logical resource pool, intelligent fiber optic patching robot, and fiber core set is selected, there is a possibility that a feasible path cannot be found from this fiber core set. If this occurs, the logical resource pools are expanded downwards in descending order according to the logical resource pool priority queue. The logical resource pool, intelligent fiber optic patching robot, and fiber core subsets are re-determined based on the intelligent fiber optic patching robot queue priority until a feasible path is found or all logical resource pools are traversed.

[0058] In one embodiment, the A-star algorithm is used to calculate the optional path set A f When a service request is received and the starting and ending nodes in the AZ are determined, the A-star algorithm is used to calculate possible fiber core paths from the starting point to the end point based on the network topology. The A-star algorithm uses a heuristic search based on the breadth-first search algorithm to estimate the cost from node n to the target node. When selecting the next node to expand at each step, the A-star algorithm prioritizes the node with the smallest f(n) value.

[0059] f(n)=φ1×g(n)+φ2×h(n)+φ3×(1-RH n )

[0060] Where g(n) is the actual cost from the starting node to node n, represented by the normalized sum of the physical lengths of all optical fiber lines from the starting node to node n; h(n) is the estimated cost from node n to the target node, represented by the normalized value of the shortest number of hops from node n to the target node; 1-RH n is the health loss cost (RH) of the intelligent fiber optic wiring robot corresponding to node n n is the normalized health score of the robot, ranging from 0 to 1); φ1, φ2, and φ3 are weights, the sum of which is 1. This allows for link calculations that consider not only the effect of optical cable length on optical signal attenuation, but also the connection relationships between subsequent nodes and the health of the currently selected intelligent fiber optic patching robot, minimizing the impact of the robot's patching on optical power loss.

[0061] See the instructions attached Figure 3 When matching reinforcement learning paths, a path is selected from the set of feasible paths as an action to be executed, and the immediate reward and the expected value of the cumulative reward are calculated based on the execution result, including the following steps:

[0062] S301. Select a path from the set of feasible paths using a greedy strategy as an action to be executed; if the generated random number is less than the set exploration rate, a path is randomly selected from the set of feasible paths as an action to be executed; if the generated random number is greater than the set exploration rate, a path with the largest expected cumulative reward is selected from the set of feasible paths as an action to be executed;

[0063] S302. Calculate the immediate reward after the action is executed based on the set reward strategy, and calculate the expected cumulative reward after the action is executed based on the immediate reward, the set learning rate and discount factor, the expected cumulative reward when the action is not executed, and the maximum expected cumulative reward among all feasible paths.

[0064] Specifically, for each path in the set of feasible paths, its value is evaluated based on the current Q value, that is, learning is done by maintaining a Q value table to record the long-term cumulative reward expectation of performing a certain action in a certain state. In the fiber optic wiring scenario, each row of the Q value table corresponds to the real-time state s at a certain time t, and each column corresponds to a possible action (i.e., the fiber core connection path, calculated by the A-star algorithm). The evaluation value Q′ (s t ,a t ), its update formula is:

[0065] Q′(s t ,at )=Q(s t ,a t )+α[r t +γmax a′ Q(s t+1 ,a′)-Q(s t ,a t )]

[0066] Among them, Q(s t ,a t ) is in state s t The fiber core connection path action a is not executed t The Q value stored at the time, Q′(s t ,a t ) is to execute the fiber core connection path action a t The Q value after s indicates that t The expected long-term cumulative reward for choosing this path; α represents the learning rate, which ranges from 0 to 1. It controls the degree to which new information is learned each time the Q value is updated. The closer α is to 1, the greater the influence of the newly acquired reward information on the update of the Q value; the closer α is to 0, the more the update of the Q value depends on previous experience; r t To perform action a t The immediate reward obtained after the action is used to measure the direct effect of the action; the γ discount factor, which ranges from [0,1]. It is used to measure the importance of future rewards. The closer γ is to 1, the more emphasis is placed on future rewards, and the long-term impact of the current action on subsequent states and rewards is considered; the closer γ is to 0, the more attention is paid to the immediate reward; s t+1 Indicates execution of action a t The new state to which the action is transferred reflects the change in state caused by the execution of the action; max a′ Q(s t+1 ,a′) represents the new state s t+1 The maximum Q value among all feasible core connection paths a′ represents the maximum expected long-term cumulative reward from selecting the optimal path in the new state. Initially, the Q value for each path in the set of feasible paths is typically initialized to a default value, such as 0. This is because sufficient experience has not yet been accumulated, and the value of each path has not been clearly judged.

[0067] Among them, the state s t It can be represented by a multi-dimensional vector, s t =[T,W,H score,M], T is the network topology information, which is used to judge the physical connectivity of the network, including the connection relationship between nodes, link length and optical power attenuation. It is composed of three N×N connection relationships (Connectivity), link length (LinkLength), and optical power attenuation (Attenuation) to form a three-dimensional matrix; W represents the dynamic weight of each logical resource pool, which is used to guide path search and is a one-dimensional array. H score This represents the predicted health value of each intelligent fiber-optic patching robot and is a one-dimensional array. M represents the fiber core occupancy status. Fiber cores are numbered according to the "resource pool-robot-fiber core" formula and marked with their occupancy status. This allows for quick location of available cores within the resource pool and is a two-dimensional array.

[0068] In the Q-value calculation formula for each path, action a represents the selection of a feasible fiber connection path from the source node to the destination node, that is, the selection of a specific fiber core combination to establish an optical path connection. For example, selecting fiber core 1 from node A, passing through fiber core 5 at intermediate node B, and finally connecting to fiber core 8 at destination node C is an action. In a certain state, an action (i.e., a path) is selected from the set of feasible paths using the following greedy strategy: a path is randomly selected from the set of feasible paths with probability ε as the action; and the path with the highest Q-value is selected as the action with probability 1-ε:

[0069]

[0070] Among them, a t is the action selected at time t; ε is the exploration rate, which ranges from [0, 1]. In the early stages of learning, ε is set to a larger value (such as 0.3) so that there are more opportunities to explore different actions; as learning progresses, ε gradually decreases, and the action currently considered optimal is more likely to be selected; A f It is a set of feasible paths that meet the optical power index requirements calculated by the A-star algorithm based on the network topology and service requests (starting and ending nodes, optical power requirements, etc.). In complex optical fiber networks, the A-star algorithm uses a heuristic function to estimate the cost from the current node to the target node, thereby quickly screening out feasible paths; Q′(s t ,a) is in state s t The Q value of executing action a under state s t The expected long-term cumulative reward of selecting action a.

[0071] Assume the current state is s t , path a is selected from the set of feasible paths t After executing the path selection action, the instant reward r is obtained based on whether the connection is successfully established and whether the optical power meets the requirements. tThen, the new Q value is calculated based on the Q value update formula. For example, if the connection is successfully established and the optical power meets the requirements, r t is positive (such as +1); if the connection fails or the optical power does not meet the standard, r t Negative (such as -1). α is the learning rate, which controls the influence of new reward information on the Q value update; γ is the discount factor, which measures the importance of future rewards. max a′ Q(s t+1 ,a′) represents the new state s t+1 The maximum Q value among all possible actions (i.e., all feasible fiber core connection paths) is determined by continuous iteration. Through continuous updates, the Q value gradually reflects the value of each path in different states. As experience accumulates, the master node evaluates the value of paths based on the updated Q value and prioritizes paths with high Q values ​​to achieve better decision-making results.

[0072] Among them, the instant reward after the action is executed is calculated based on the set reward strategy, and the reward strategy can be set according to one or more indicators of connection quality, fiber core health, business goals, and path efficiency. For example, according to the selected action, a fiber jumper instruction is issued to the designated intelligent fiber optic wiring robot, and the intelligent fiber optic wiring robot performs the fiber core connection operation. After the connection is completed, if a connection that meets the optical power index requirements is successfully established, an instant reward of +1 is given; if the connection fails, an instant reward of -1 is given; if the connection is successfully established, but exceeds the optical power index threshold, a smaller negative reward (such as -0.5) is given. The specific setting is based on the actual application, and this application does not limit or fix it.

[0073] The above process of acquiring state information, selecting actions, executing actions, calculating rewards, and updating policies is repeated. As the number of iterations increases, the policy is continuously learned and optimized, gradually finding a more optimal fiber core path selection method. By combining the A-star algorithm with Q-value reinforcement learning, the action selection phase can more effectively utilize network topology information to find the optimal fiber core path for service requests. This overcomes the limitations of traditional path selection methods, eliminating blind searches or reliance on a single factor. Furthermore, the learning capabilities of reinforcement learning are leveraged to continuously optimize policies, improve SLA (Service-Level Agreement) compliance and resource utilization, and adapt to dynamic changes in network status.

[0074] It can be seen that the application provides an intelligent fiber optic wiring scheduling method based on reinforcement learning decision-making, which abstracts the physical fiber core into a logical resource pool and uses dynamic weight partition scheduling to achieve optimal resource allocation. It breaks through the physical limitations of traditional resource allocation and is no longer limited to the fiber core range managed by a single intelligent fiber optic wiring robot. By calculating dynamic weights for each logical resource pool based on real-time load and fiber core health, business requests can be accurately allocated to the optimal robot management area, achieving robot load balancing and overload avoidance. In addition, through reinforcement learning technology, business needs and resource status are matched in real time, and the optimal fiber core path is dynamically generated. For application scenarios with high requirements for service activation and recovery speed, it can quickly find the fiber core path that meets the business requirements, improve business response efficiency, and reduce business waiting time.

[0075] Based on the same inventive concept, an embodiment of the present application also provides an intelligent fiber optic wiring scheduling device based on reinforcement learning decision-making. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned intelligent fiber optic wiring scheduling method based on reinforcement learning decision-making in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0076] As the instruction manual Figure 4 As shown, the embodiment of the present application also provides an intelligent optical fiber wiring scheduling device based on reinforcement learning decision-making, which is applied to a cluster composed of multiple intelligent optical fiber wiring robots, and the device includes:

[0077] Construction module 401 is used to abstract the fiber cores connected to different intelligent fiber distribution robots in the cluster into a logical resource pool;

[0078] A determination module 402 is configured to determine the priority of the logical resource pool and the intelligent fiber optic wiring robot contained therein;

[0079] The path matching module 403 is used to determine a scheduling path that meets business requirements based on a reinforcement learning algorithm. First, a search algorithm is used to screen a set of feasible paths that meet business requirements in a logical resource pool determined based on the priority. Then, the current state of the optical fiber network is obtained, and a path is selected from the set of feasible paths as an action to be executed. The immediate reward and the expected value of the cumulative reward are calculated based on the execution result. Finally, based on the calculated expected value of the cumulative reward, the path is iteratively optimized to determine a scheduling path that meets business requirements.

[0080] In one embodiment, the construction module 401 abstracts the fiber cores connected to different intelligent fiber optic distribution robots in the cluster into logical resource pools, including: configuring logical identifiers for the fiber core range managed by each intelligent fiber optic distribution robot; and calculating the dynamic weight of each logical resource pool based on the load condition and average health of the fiber core.

[0081] In one embodiment, the determination module 402 determines the priority of the logical resource pool and the intelligent fiber optic distribution robots contained therein, including: determining the priority of the logical resource pool according to the size of the dynamic weight; determining the priority of each intelligent fiber optic distribution robot in the logical resource pool according to the health of the intelligent fiber optic distribution robot and the number of idle fiber cores.

[0082] In one embodiment, the path matching module 403 uses a search algorithm to screen a set of feasible paths that meet business requirements in the logical resource pool determined based on the priority, including: selecting a set number of logical resource pools and the intelligent fiber optic distribution robots contained therein in order of priority to obtain a managed fiber core set; using the A-star algorithm based on network topology information to calculate a set of feasible paths from the source node to the destination node from the fiber core set; judging whether there is a path that meets business requirements in the feasible path set, and if so, using the feasible path set as the action space of reinforcement learning; if not, expanding the number of selected logical resource pools and the intelligent fiber optic distribution robots contained therein in order of priority, and recalculating the feasible path set until a path that meets business requirements is found.

[0083] In one embodiment, the path matching module 403 selects a path from the set of feasible paths as an action to be executed, and calculates an immediate reward and an expected cumulative reward based on the execution result, including: using a greedy strategy to select a path from the set of feasible paths as an action to be executed; wherein, if the generated random number is less than a set exploration rate, a path is randomly selected from the set of feasible paths as an action to be executed; if the generated random number is greater than the set exploration rate, a path with the largest expected cumulative reward is selected from the set of feasible paths as an action to be executed; calculating an immediate reward after the action is executed based on a set reward strategy, and calculating an expected cumulative reward after the action is executed based on the immediate reward, a set learning rate and discount factor, the expected cumulative reward when the action is not executed, and the maximum expected cumulative reward among all feasible paths. The reward strategy is set based on one or more indicators of connection quality, fiber core health, service objectives, and path efficiency.

[0084] The intelligent fiber optic patching scheduling device based on reinforcement learning decision-making described in this application is applied to a cluster composed of multiple intelligent fiber optic patching robots. A construction module abstracts the fiber cores of different intelligent fiber optic patching robots in the cluster into a logical resource pool. A determination module determines the priority of the logical resource pool and the intelligent fiber optic patching robots it contains. A path matching module determines a scheduling path that meets business requirements based on a reinforcement learning algorithm. A search algorithm is first used to screen a set of feasible paths that meet business requirements in the logical resource pool determined based on the priority. Then, the current state of the optical fiber network is obtained, and a path is selected from the set of feasible paths as an action to be executed. Based on the execution result, an immediate reward and an expected cumulative reward are calculated. Finally, based on the calculated expected cumulative reward, the path is iteratively optimized to determine a scheduling path that meets business requirements. Thus, by combining reinforcement learning with dynamic weighted scheduling, business needs and resource status are matched in real time, and the optimal fiber core path is dynamically generated. For application scenarios with high service provisioning and recovery speed requirements, the device can quickly find a fiber core path that meets business requirements, improve business response efficiency, and reduce business waiting time.

[0085] Based on the same concept of the present invention, as shown in the attached specification Figure 5 As shown, an embodiment of the present application provides a structure of an electronic device 500, which includes: at least one processor 501, at least one network interface 504 or other user interface 503, a memory 505, and at least one communication bus 502. The communication bus 502 is used to achieve connection and communication between these components. The electronic device 500 optionally includes a user interface 503, including a display (for example, a touch screen, LCD, CRT, holographic imaging (Holographic) or projection (Projector), etc.), a keyboard or a pointing device (for example, a mouse, trackball (trackball), touchpad or touch screen, etc.).

[0086] The memory 505 may include a read-only memory and a random access memory, and provides instructions and data to the processor 501. A portion of the memory 505 may also include a non-volatile random access memory (NVRAM).

[0087] In some embodiments, the memory 505 stores the following elements, executable modules, or data structures, or a subset or extended set thereof:

[0088] Operating system 5051, including various system programs for implementing various basic services and processing hardware-based tasks;

[0089] The application module 5052 includes various application programs, such as a launcher, a media player, a browser, etc., which are used to implement various application services.

[0090] In the embodiment of the present application, by calling the program or instructions stored in the memory 505, the processor 501 is used to execute the steps of an intelligent fiber optic wiring scheduling method based on reinforcement learning decision-making.

[0091] The present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program executes steps in an intelligent optical fiber distribution scheduling method based on reinforcement learning decision-making.

[0092] Specifically, the storage medium can be a general storage medium, such as a mobile disk, hard disk, etc. When the computer program on the storage medium is run, it can use reinforcement learning technology to match business needs and resource status in real time, dynamically generate the optimal fiber core path, and improve business response efficiency.

[0093] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0094] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0095] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0096] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0097] Finally, it should be noted that the above embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The scope of protection of the present application is not limited thereto. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above embodiments within the technical scope disclosed in the present application, or replace some of the technical features therein with equivalents. However, these modifications, changes, or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application. They should all be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. An intelligent optical fiber distribution scheduling method based on reinforcement learning decision making, characterized in that: Applied to a cluster consisting of multiple intelligent fiber optic wiring robots, the method includes the following steps: Abstract the fiber cores connected to different intelligent fiber distribution robots in the cluster into a logical resource pool; Determining the priority of the logical resource pool and the intelligent fiber optic wiring robot contained therein; A scheduling path that meets business requirements is determined based on a reinforcement learning algorithm. Specifically, a search algorithm is first used to screen a set of feasible paths that meet business requirements in a logical resource pool determined based on the priority. Then, the current state of the optical fiber network is obtained, and a path is selected from the set of feasible paths as an action to be executed. An immediate reward and an expected value of a cumulative reward are calculated based on the execution result. Finally, based on the calculated expected value of the cumulative reward, the path is iteratively optimized to determine a scheduling path that meets business requirements.

2. The intelligent optical fiber distribution scheduling method based on reinforcement learning decision-making according to claim 1 is characterized in that: The process of abstracting the fiber cores connected to different intelligent fiber distribution robots in the cluster into a logical resource pool includes the following steps: Configure logical identification for the fiber core range managed by each intelligent fiber optic wiring robot; The dynamic weight of each logical resource pool is calculated based on its load and average fiber core health.

3. The intelligent optical fiber distribution scheduling method based on reinforcement learning decision-making according to claim 2 is characterized in that: Determining the priority of the logical resource pool and the intelligent optical fiber distribution robot contained therein comprises the following steps: Determining the priority of the logical resource pool according to the size of the dynamic weight; The priority of each intelligent fiber optic patching robot in the logical resource pool is determined based on the health status and number of idle fiber cores of the intelligent fiber optic patching robot.

4. The intelligent optical fiber distribution scheduling method based on reinforcement learning decision-making according to claim 3 is characterized in that: The method of using a search algorithm to screen a set of feasible paths that meet business requirements in a logical resource pool determined based on the priority includes the following steps: Selecting a set number of logical resource pools and the intelligent fiber optic wiring robots contained therein in order of priority to obtain a managed fiber core set; Calculating a set of feasible paths from a source node to a destination node from the set of fiber cores using an A-star algorithm based on network topology information; Determine whether there is a path in the feasible path set that meets the business requirements. If so, use the feasible path set as the action space for reinforcement learning; if not, expand the selected logical resource pool and the number of intelligent fiber optic distribution robots it contains in order of priority, and recalculate the feasible path set until a path that meets the business requirements is found.

5. The intelligent optical fiber distribution scheduling method based on reinforcement learning decision-making according to claim 4 is characterized in that: in, The status of the optical fiber network includes network topology information, dynamic weights of the logical resource pool, and the health of the intelligent optical fiber distribution robot; the network topology information includes the connection relationship between nodes, link length, and optical power attenuation.

6. The intelligent optical fiber distribution scheduling method based on reinforcement learning decision-making according to claim 5, characterized in that: The step of selecting a path from the set of feasible paths as an action to be executed, and calculating an immediate reward and an expected value of a cumulative reward according to the execution result, comprises the following steps: A greedy strategy is used to select a path from the set of feasible paths as an action to be executed; if the generated random number is less than the set exploration rate, a path is randomly selected from the set of feasible paths as an action to be executed; if the generated random number is greater than the set exploration rate, the path with the largest expected cumulative reward is selected from the set of feasible paths as an action to be executed; The immediate reward after the action is executed is calculated based on the set reward strategy, and the expected cumulative reward after the action is executed is calculated based on the immediate reward, the set learning rate and discount factor, the expected cumulative reward when the action is not executed, and the maximum expected cumulative reward among all feasible paths.

7. The intelligent optical fiber distribution scheduling method based on reinforcement learning decision-making according to claim 6, characterized in that: in, The reward strategy is set according to one or more indicators of connection quality, fiber core health, service goals, and path efficiency.

8. An intelligent optical fiber wiring scheduling device based on reinforcement learning decision-making, characterized in that: Applied to a cluster consisting of multiple intelligent fiber optic wiring robots, the device includes: A building block for abstracting the fiber cores connected to different intelligent fiber distribution robots in the cluster into a logical resource pool; A determination module, configured to determine the priority of the logical resource pool and the intelligent fiber optic wiring robot contained therein; A path matching module is used to determine a scheduling path that meets business requirements based on a reinforcement learning algorithm. A search algorithm is first used to screen a set of feasible paths that meet business requirements from a logical resource pool determined based on the priority level. The current state of the optical fiber network is then obtained, and a path is selected from the set of feasible paths as an action to be executed. An immediate reward and an expected cumulative reward are calculated based on the execution result. Finally, based on the calculated expected cumulative reward, the path is iteratively optimized to determine a scheduling path that meets business requirements.

9. An electronic device, characterized in that: include: A processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate via the bus. When the machine-readable instructions are executed by the processor, the steps of the intelligent optical fiber distribution scheduling method based on reinforcement learning decision-making are performed as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the intelligent optical fiber distribution scheduling method based on reinforcement learning decision-making as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Optimal optical fiber path selection method and device

    CN112511230A

  • A star algorithm improvement method, system and device based on dynamic weight and medium

    CN114527788A

  • Optical fiber distribution selection and scheduling method of cloud computing data center

    CN115882949A

  • Intelligent network path optimization method and system based on deep reinforcement learning

    CN116527567A

  • Task execution method and device of optical fiber wiring robot, electronic equipment and medium

    CN116582180A