Network optimization method and network system
By obtaining multi-parameter state in the optical-wireless converged network and optimizing network resource allocation in combination with multiple evaluation indicators, the problem of limited optimization effect of single parameter in the existing technology is solved, and more efficient network performance improvement is achieved. It is suitable for scenarios with high latency requirements such as factory AGV and anti-optical link jitter scenarios such as hospital mobile surgical vehicles.
Patent Information
- Application Number
- CN202510822362.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-08-12
AI Technical Summary
In existing optical-wireless converged networks, network resource adjustment is often based on a single parameter, resulting in limited optimization effects and prone to non-optimal adjustment solutions, such as ‘strong signal but high delay’ or ‘low load but high interference’.
By obtaining the status parameters of the fiber optic network and wireless network, and combining multiple evaluation indicators, the target control strategy is determined to optimize network resource allocation, including the connection relationship of terminal devices and the adjustment of network bandwidth resources.
It improves network performance, solves the problems of insufficient bandwidth and network instability of terminal devices, and meets high latency requirements such as real-time control of factory AGV and anti-optical link jitter scenarios such as hospital mobile surgical vehicle control.
Smart Images

Figure CN120475466A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of wireless network technology, and in particular to a network optimization method and a network system. Background Art
[0002] The optical-wireless converged network architecture is a network architecture that integrates optical fiber and wireless networks. In this network architecture, wireless access points can access the core network through optical links, and each terminal device can connect to the wireless network provided by the wireless access points. This network architecture has the advantages of high bandwidth of optical links and the ease of use of wireless networks.
[0003] To improve network performance, converged optical-wireless networks often require adjustments to network resources, such as controlling the handoff of end devices from one wireless access point to another. Existing methods typically optimize based on a single network parameter, such as wireless signal strength or access point load. This approach offers limited optimization results and can easily lead to suboptimal adjustments, such as "strong signal but high latency" or "low load but high interference." Summary of the Invention
[0004] To solve the above problems, this application discloses the following technical solutions:
[0005] A first aspect of the present application provides a network optimization method, comprising:
[0006] Obtaining state parameters of a target network system consisting of a fiber optic network and a wireless network, the state parameters including wireless parameters reflecting a network environment of the wireless network, optical layer parameters reflecting a network environment of the fiber optic network, and terminal parameters reflecting requirements of terminal devices within a future preset time period;
[0007] With the goal of optimizing the long-term cumulative reward corresponding to the control strategy, a target control strategy is determined according to the state parameter, wherein the long-term cumulative reward is determined by a plurality of evaluation indicators, wherein the plurality of evaluation indicators include any multiple of a first indicator reflecting the success rate of the control strategy, a second indicator reflecting the resources consumed by the control strategy, and a third indicator reflecting the effect achieved by the control strategy;
[0008] The network resource allocation of the target network system is optimized according to the target control strategy.
[0009] Optionally, at least one of the following is also included:
[0010] In response to a preset control cycle, executing the step of obtaining the state parameters of the target network system consisting of the optical fiber network and the wireless network;
[0011] In response to detecting a target triggering event, executing the step of obtaining the state parameters of the target network system consisting of the optical fiber network and the wireless network;
[0012] The target trigger event includes an optical layer parameter satisfying a first trigger condition and / or a wireless parameter satisfying a second trigger condition.
[0013] Optionally, also include:
[0014] Fault risk information of the optical fiber network is determined according to the optical layer parameters, where the fault risk information at least represents a probability of a potential fault existing in the optical fiber network.
[0015] Optionally, the target control strategy includes multiple adjustment actions and an action probability corresponding to each adjustment action;
[0016] Optimizing the network resource allocation of the target network system according to the target control strategy includes:
[0017] determining a target adjustment action in the target control strategy based on the corresponding action probability, and executing the target adjustment action to adjust the network resource allocation of the target network system;
[0018] The multiple adjustment actions include actions for adjusting the connection relationship between the terminal device and the target network system, and / or actions for adjusting bandwidth resources of the optical fiber network and the wireless network.
[0019] Optionally, the performing the target adjustment action to adjust the network resource allocation of the target network system includes at least one of the following:
[0020] Controlling a first terminal device to switch from a first wireless access point to a second wireless access point, where the first wireless access point is the wireless access point currently accessed by the first terminal device;
[0021] Migrating part of the network services of the first optical network unit connected to the third wireless access point on the current wavelength channel to another idle wavelength channel, wherein the third wireless access point and the second wireless access point are the same or different wireless access points;
[0022] Adjust the optical layer bandwidth resources allocated to each connected wireless access point by the second optical network unit, where the second optical network unit and the first optical network unit are the same or different optical network units.
[0023] Optionally, the long-term cumulative reward is determined by multiple evaluation indicators, including:
[0024] The long-term cumulative reward is determined based on the immediate reward value, which is obtained by processing multiple evaluation indicators according to a preset reward function;
[0025] After optimizing the network resource allocation of the target network system according to the target control strategy, the method further includes:
[0026] Obtain optimization effect indicators;
[0027] Adjusting the reward function according to the optimization effect indicator;
[0028] The optimization effect indicator includes at least one of switching time and service interruption duration.
[0029] Optionally, the step of determining a target control strategy based on the state parameters with the goal of optimizing the long-term cumulative reward corresponding to the control strategy includes:
[0030] Processing the state parameters according to a pre-built strategy network to obtain a current control strategy;
[0031] Processing the current control strategy according to a pre-built value network to obtain a current long-term cumulative reward corresponding to the current control strategy;
[0032] If the current long-term cumulative reward does not satisfy the target optimization condition, updating the current control strategy in the strategy network based on the state parameter and the current long-term cumulative reward until a current control strategy in which the current long-term cumulative reward satisfies the target optimization condition is obtained;
[0033] The current control strategy whose current long-term cumulative reward satisfies the target optimization condition is determined as the target control strategy.
[0034] Optionally, the optical layer parameter includes at least one of optical fiber link delay, optical power fluctuation, wavelength occupancy, available wavelength resources, optical fiber link delay standard deviation, bandwidth utilization and bit error rate;
[0035] And / or, the wireless parameter includes at least one of a channel signal-to-noise ratio, co-channel interference intensity, number of user accesses, and wireless access point load rate.
[0036] Optionally, the terminal parameters include at least one of an expected movement trajectory parameter and a service traffic type of the terminal device, and the expected movement trajectory parameter is determined based on a historical location parameter, a movement speed parameter, and a movement acceleration parameter of the terminal device.
[0037] A second aspect of the present application provides a network system including a wireless network and an optical fiber network;
[0038] The wireless network includes a terminal device and a wireless access point for accessing the terminal device;
[0039] The optical fiber network includes an optical line terminal and an optical network unit connected by optical fibers, and the optical network unit and the wireless access point are connected by optical fibers;
[0040] The optical network unit is used for:
[0041] Obtaining state parameters of a target network system consisting of a fiber optic network and a wireless network, the state parameters including wireless parameters reflecting a network environment of the wireless network, optical layer parameters reflecting a network environment of the fiber optic network, and terminal parameters reflecting terminal demand within a future preset time period;
[0042] With the goal of optimizing the long-term cumulative reward corresponding to the control strategy, a target control strategy is determined according to the state parameter, wherein the long-term cumulative reward is determined by a plurality of evaluation indicators, wherein the plurality of evaluation indicators include any multiple of a first indicator reflecting the success rate of the control strategy, a second indicator reflecting the resources consumed by the control strategy, and a third indicator reflecting the effect achieved by the control strategy;
[0043] The network resource allocation of the target network system is optimized according to the target control strategy.
[0044] This solution obtains the wireless parameters of the wireless network in the target network system, the optical layer parameters of the fiber network, and the terminal parameters of the terminal devices. Based on these multiple parameters, it determines the target control strategy. This strategy is then determined by optimizing the long-term cumulative reward across multiple evaluation metrics. Therefore, compared to existing solutions that analyze only a single parameter and adjust based on optimizing a single metric, this solution can better optimize network performance and address issues such as insufficient bandwidth and network instability in the target network system's terminal devices.
[0045] Based on the above advantages, this solution can be applied to scenarios with high latency requirements, such as real-time control of factory AGVs with a latency requirement of less than 10 milliseconds, and scenarios requiring strong resistance to optical link jitter, such as control of mobile operating carts in hospitals. This solves the pain point of existing solutions where strong wireless signals but congested optical links lead to excessive latency. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0047] Figure 1This is a flow chart of a network optimization method provided by an embodiment of the present application;
[0048] Figure 2 This is a schematic diagram of the architecture of a target network system provided by an embodiment of the present application;
[0049] Figure 3 This is a flow chart of a method for determining a target control strategy provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0051] This embodiment provides a network optimization method, see Figure 1 , is a flowchart of the method, which may include the following steps.
[0052] S101, obtain state parameters of a target network system consisting of a fiber optic network and a wireless network connection, the state parameters including wireless parameters reflecting the network environment of the wireless network, optical layer parameters reflecting the network environment of the fiber optic network, and terminal parameters reflecting the needs of terminal devices within a preset time period in the future.
[0053] S102, with the goal of optimizing the long-term cumulative reward corresponding to the control strategy, determine the target control strategy based on the state parameters. The long-term cumulative reward is determined by multiple evaluation indicators, and the multiple evaluation indicators include any multiple of a first indicator reflecting the success rate of the control strategy, a second indicator reflecting the resources consumed by the control strategy, and a third indicator reflecting the effect achieved by the control strategy.
[0054] S103: Optimize the network resource allocation of the target network system according to the target control strategy.
[0055] This solution obtains the wireless parameters of the wireless network in the target network system, the optical layer parameters of the fiber network, and the terminal parameters of the terminal devices. Based on these multiple parameters, it determines the target control strategy. This strategy is then determined by optimizing the long-term cumulative reward across multiple evaluation metrics. Therefore, compared to existing solutions that analyze only a single parameter and adjust based on optimizing a single metric, this solution can better optimize network performance and address issues such as insufficient bandwidth and network instability in the target network system's terminal devices.
[0056] For reference, compared to a solution that uses the A3 event for receiving signal strength indication (RSSI) to switch terminals, this solution reduces ping-pong handovers by 73% in high-density scenarios, reducing service interruption duration from 120 milliseconds to 15ms, meeting the latency requirement of no more than 20ms in industrial control scenarios. High-density scenarios generally refer to scenarios where each wireless access point has an average of 200 or more terminals connected.
[0057] The network optimization method of this embodiment can be applied to a target network system, which can be any network system composed of an interconnected optical fiber network and a wireless network, that is, an optical-wireless network system.
[0058] From a network architecture perspective, the target network system can include modules such as optical line terminals, optical network units, wireless access points, and terminal devices. The optical line terminals and optical network units constitute the fiber optic network portion of the target network system, while the wireless access points and terminal devices constitute the wireless network portion of the target network system.
[0059] A target network system may include one or more optical line terminals (OLTs). For example, see Figure 2 Figure 2 shows an architecture diagram of a target network system, which includes an optical line terminal (OLT). An OLT is typically located in a central computer room or at a telecom operator's central office. It initiates and manages fiber-optic communications, providing data and communication services to the optical network unit (ONU).
[0060] An optical line terminal can be connected to one or more optical network units (ONUs) through optical fibers (also called optical links). Figure 2 For example, the optical line terminal is connected to optical network unit 1 and optical network unit 2 via optical fiber. Generally, the optical network unit can be located at the user end, such as a home, office or other place where network access is required, and serves as an interface device for wireless access points to access the optical fiber network.
[0061] Optical network units (ONUs) can have built-in embedded neural network processors (NPUs), which can integrate deep reinforcement learning (DRL) inference engines. This NPU only requires less than 15% of the ONU's computing power to implement the method of this embodiment. As edge nodes, the ONUs can prioritize local decision-making to implement the aforementioned optimization method and regularly synchronize model parameters to the cloud via optical links (for example, at dawn each day), achieving "edge-first, cloud-assisted" collaborative optimization.
[0062] An optical network unit can be connected to one or more wireless access points (AP) through optical fiber. Figure 2 For example, optical network unit 1 can be connected to wireless access point 1, wireless access point 2, and wireless access point 3 via optical fiber. Each wireless access point can provide a wireless network, such as a WiFi network, to an area within a certain range. The wireless access point can be any device that can provide a wireless network and connect to an optical network unit, and its type is not limited. As an example, the wireless access point used in this embodiment can be a Wi-Fi 7 AP / 5G small base station device.
[0063] When any terminal device is within the coverage of a wireless access point, the terminal device can access the wireless access point, and thus access the core network through the wireless access point and the back-end optical network unit, optical line terminal and other equipment. Figure 2 For example, there may be terminal device 1 and terminal device 2 within the coverage of wireless access point 1, and both terminal device 1 and terminal device 2 access the core network through the wireless network provided by wireless access point 1.
[0064] The optimization method of this embodiment can be performed by edge nodes of the target network system. Here, edge nodes can refer to optical network units (ONUs) in the target network system. In other words, each ONU can perform the optimization method of this embodiment to optimize and adjust the network resource allocation of its connected wireless access points and their terminal devices.
[0065] The terminal device of this embodiment can be any electronic device that can access a wireless network and is movable or portable. The type of terminal device can be different depending on the application scenario of the target network system and is not limited to this.
[0066] As some examples, when used in industrial scenarios, the terminal device can be an automated guided vehicle (AGV); when used in campus or large supermarket scenarios, the terminal device can be a consumer electronic product held or worn by the user, such as a smartphone or smart watch; when used in large hospital scenarios, the terminal device can be various networked medical devices, such as operating tables, monitors, ventilators, etc.
[0067] The following combination Figure 1 The specific implementation process of the method of this embodiment is described.
[0068] In step S101, the optical layer parameters may include only the physical layer parameters of the optical fiber network, or only the transport layer parameters of the optical fiber network, or both the physical layer parameters and the transport layer parameters of the optical fiber network.
[0069] The physical layer parameters of the optical fiber network may include any one or more of the optical fiber link delay, optical power fluctuation and wavelength occupancy. These physical layer parameters can be obtained through the ONU or OLT detection of the optical fiber network. The detection method can be referred to the relevant existing technology and will not be repeated here.
[0070] The transport layer parameters of the optical fiber network may include either or both of bandwidth utilization and bit error rate, and may also include other optional transport layer parameters. In this embodiment, the edge node may parse the transmitted signal using a relevant transmission protocol to obtain the transport layer parameters. The transmission protocol used may be the IEEE 802.3ah protocol or other transmission protocols applicable to optical fiber networks.
[0071] In summary, the optical layer parameters obtained in S101 may include any one or more of optical fiber link delay, optical power fluctuation, wavelength occupancy, available wavelength resources, optical fiber link delay standard deviation, bandwidth utilization and bit error rate.
[0072] Wireless parameters may include any one or more of the following: channel signal-to-noise ratio (SNR), co-channel interference intensity, number of connected users, and wireless access point load. These parameters, such as channel SNR, co-channel interference intensity, and number of connected users, can be monitored by wireless access points. Specifically, a channel probing module for channel quality monitoring can be pre-deployed in each wireless access point. This module monitors the corresponding wireless parameters in real time and provides feedback to the edge node to which the wireless access point is connected. Wireless parameters may also include base station load.
[0073] Optionally, the terminal parameters may include any one or a combination of the terminal device's expected movement trajectory parameters and the service traffic type. The expected movement trajectory parameters are determined based on the terminal device's historical location parameters, movement speed parameters, and movement acceleration parameters. The terminal device's historical location parameters, movement speed parameters, and movement acceleration parameters may be monitored by the wireless access point to which the terminal device accesses via a pre-deployed monitoring module, or the terminal device may monitor and report these to the wireless access point via its own monitoring module.
[0074] The service traffic type may be determined by analyzing service data transmitted between the wireless access point and the terminal device based on a relevant transmission protocol. The transmission protocol used may be the IEEE 802.11k / v protocol, or other transmission protocols, without limitation.
[0075] The wireless access point can feed back the historical location parameters, movement speed parameters, and movement acceleration parameters of a terminal device to the edge node. The edge node can then process these parameters of the terminal device based on the Kalman filter algorithm or other algorithms that can predict movement trajectories, thereby predicting the expected movement trajectory parameters of the terminal device within a specific time period after the current moment. The specific time period here can be the next 500 milliseconds (ms) from the current moment, or the next 2 seconds from the current moment. The specific setting can be based on actual needs and the performance of the algorithm used, and is not limited.
[0076] The expected movement trajectory parameters of a terminal device can be in any form and can represent the parameters of the movement trajectory of the terminal device within the corresponding time. As some examples, the expected movement trajectory parameters may include coordinate parameters of several moments in the corresponding time, each coordinate parameter representing the expected position of the terminal device at the corresponding moment. Multiple coordinate parameters can represent the movement trajectory of the terminal device.
[0077] Exemplarily, 500 milliseconds may be divided into multiple moments at intervals of 10 milliseconds, and the expected movement trajectory parameters of a terminal device in the next 500 milliseconds may include the coordinate parameters of the terminal device at each moment.
[0078] During the operation of the target network system, if the edge node determines that a specific execution condition is met at a certain moment, step S101 can be executed once to obtain a state parameter at the current moment to determine the target control strategy based on the state parameter.
[0079] Among the status parameters obtained by the edge node, the optical layer parameters may include the optical layer parameters between the edge node itself and the connected wireless access point, the optical layer parameters between the edge node and the optical line terminal connected to itself, or both; the wireless parameters may include the wireless parameters between each wireless access point under the edge node and each terminal device accessing the wireless access point; and the terminal parameters may include the terminal parameters of each terminal device connected to each wireless access point under the edge node.
[0080] by Figure 2 For example, as an edge node, the status parameters obtained by the optical network unit 1 may include optical layer parameters between the optical network unit 1 and the optical line terminal, optical layer parameters between the optical network unit 1 and wireless access points 1 to 3, wireless parameters between the wireless access point 1 and terminal devices 1 and 2, and terminal parameters of terminal devices 1 and 2, that is, expected movement trajectory parameters of terminal device 1 and expected movement trajectory parameters of terminal device 2.
[0081] Optionally, before executing S102 according to the state parameters, the state parameters may be preprocessed first, and then S102 may be executed based on the preprocessed state parameters. This is conducive to obtaining a target control strategy that better meets the actual needs of the target network system.
[0082] The preprocessing method for the state parameters may include any one or more of the following.
[0083] Preprocessing method 1 involves time synchronization of the optical layer parameters and wireless parameters in the status parameters. For example, the IEEE 1588 Precision Time Protocol (PTP) can be used to synchronize the optical layer parameters and wireless parameters, ensuring that the time error between the two parameters is less than 1 microsecond (μs). For methods of time synchronization based on the precision time protocol, refer to related prior art and are not described here.
[0084] Preprocessing method 2 is to normalize heterogeneous data. Heterogeneous data can include optical power in optical layer parameters, or other optical layer parameters or wireless parameters. The following uses optical power as an example for explanation.
[0085] First, determine the upper and lower limits of optical power based on the physical characteristics and protocol standards of the devices in the fiber optic network (i.e., ONUs and OLTs). For example, if the minimum detectable power of the optical receivers used in the ONUs and OLTs is -28 decibel milliwatts (dBm) and the maximum input power is -3dBm, the former can be determined as the lower limit of optical power, and the latter as the upper limit of optical power. The actual measured optical power can then be calculated using the following formula to obtain the normalized optical power:
[0086] Normalized optical power = (upper optical power limit − lower optical power limit) / (actual optical power − lower optical power limit).
[0087] Preprocessing method three, the optical layer parameters, wireless parameters and terminal parameters are integrated to form a state parameter matrix. In this preprocessing method, the optical layer parameters can be combined into a row vector representing the optical resources, and each element of the vector is one of the aforementioned optical layer parameters; the wireless parameters can be combined into a row vector representing the wireless resources, and each element of the vector is one of the aforementioned wireless parameters; the terminal parameters can be combined into a row vector representing the user needs, and each element of the vector is equivalent to the expected movement trajectory parameter of a terminal device. The three row vectors are then combined to form a three-row matrix, which serves as the state parameter matrix. Alternatively, the state parameter matrix can be a 12-dimensional vector, in which the 1st to 5th dimensions are optical layer parameters, the 6th to 9th dimensions are wireless parameters, and the 10th to 12th dimensions are terminal parameters.
[0088] If the processing is performed according to the third preprocessing method, the state parameters used subsequently can be the state parameter matrix obtained by fusion.
[0089] Optionally, the specific execution condition for triggering the edge node to obtain the state parameter in the aforementioned embodiment may include at least one of the following:
[0090] Execution condition one: in response to a predetermined control cycle, executing the step of obtaining a state parameter of a target network system consisting of an optical fiber network and a wireless network;
[0091] Execution condition two: in response to detecting a target trigger event, executing the step of obtaining a state parameter of a target network system consisting of an optical fiber network and a wireless network;
[0092] The target trigger event includes an optical layer parameter satisfying a first trigger condition and / or a wireless parameter satisfying a second trigger condition.
[0093] That is, a specific execution condition may include only execution condition one, or only execution condition two, or both execution conditions one and two.
[0094] In Execution Condition 1, the control period can be set as needed, for example, to 200ms, 600ms, 1 second, etc., without limitation. Taking a control period of 200ms as an example, if the method of this embodiment is triggered according to Execution Condition 1, each edge node in the target network system can execute the optimization method of this embodiment every 200ms after startup.
[0095] In execution condition 2, the first trigger condition can be that the corresponding optical layer parameter is greater than or equal to a preset first trigger threshold. The optical layer parameter used for triggering can be optical link latency, or it can be another optical layer parameter. Taking optical link latency as an example, the first trigger condition can be that the optical link latency is greater than 50 microseconds. Here, 50 microseconds is an example of the first trigger threshold. In actual application, it can be replaced with other values, such as 100 microseconds.
[0096] The second trigger condition may be that the corresponding wireless parameter is greater than or equal to a preset second trigger threshold. The triggering wireless parameter may be wireless channel utilization, or other wireless parameters. Taking wireless channel utilization as an example, the second trigger condition may be that the wireless channel utilization is greater than 80%. 80% is just an example of the second trigger threshold, and other values, such as 70%, may be used in actual applications.
[0097] When the method of this embodiment is triggered according to execution condition 2, if the edge node finds that the optical layer parameters at the current moment meet the first trigger condition, or finds that the wireless parameters at the current moment meet the second trigger condition, or finds that the optical layer parameters at the current moment meet the first trigger condition and the wireless parameters at the current moment meet the second trigger condition, it can execute S101 and perform subsequent steps based on the obtained status parameters.
[0098] In some optional embodiments, after obtaining the state parameters, the edge node may further perform the following steps based on the optical layer parameters:
[0099] The fault risk information of the optical fiber network is determined according to the optical layer parameters. The fault risk information at least represents the probability of a potential fault existing in the optical fiber network.
[0100] In order to obtain fault risk information, a pre-built optical link health assessment model can be deployed on the edge node in advance. Each time after obtaining the status parameters, the edge node can call the optical link health assessment model to process the optical layer parameters therein to obtain the fault risk information of the optical fiber network at the current moment.
[0101] The form of fault risk information is not limited. In some embodiments, the fault risk information may include several common faults in the optical fiber network and the corresponding probability of occurrence of each fault. In some embodiments, the fault risk information may include a health parameter that reflects whether the optical fiber network as a whole can operate normally, and the type of fault that may occur when the health parameter falls below a certain threshold.
[0102] As some examples, common faults in optical fiber networks may include, but are not limited to, optical fiber aging, optical fiber breakage, optical receiver downtime, etc., without limitation.
[0103] After obtaining the fault risk information, a decision can be made based on the fault risk information whether to execute steps S102 and S103. For example, if the fault risk information indicates that the optical fiber network is less likely to fail, then steps S102 and S103 can be executed. If the fault risk information indicates that the optical fiber network is more likely to fail, then steps S102 and S103 can be temporarily skipped. In this case, the fault risk information and the collected optical layer parameters can be sent to a terminal used to maintain the optical fiber network, allowing relevant operation and maintenance personnel to perform inspections and maintenance based on the fault risk information and the collected optical layer parameters.
[0104] Among them, if the fault risk information includes several common faults in the optical fiber network and the corresponding probability of occurrence of each fault, then when each probability is less than a preset probability threshold, it can be determined that the possibility of the fault is small, and when one or several probabilities are greater than or equal to the probability threshold, it can be determined that the possibility of the fault is large; when the fault risk information includes a health parameter that reflects whether the optical fiber network as a whole can operate normally, if the health parameter is less than a preset threshold, it can be determined that the possibility of the fault is large, and if the health parameter is greater than or equal to the threshold, it can be determined that the possibility of the fault is small.
[0105] The optical link health assessment model can be a neural network model of any structure pre-trained based on sample data containing optical layer parameters, or other models that can be used for fault detection (such as classification models). The specific structure and construction method of the model can be found in the prior art regarding the structure and construction method of the fault detection model or fault identification model, which will not be repeated here.
[0106] In step S102, the edge node can use a pre-built deep reinforcement learning (DRL) model to process state parameters to obtain a target control strategy. The DRL model can include two parts: a policy network and a value network.
[0107] The policy network, also known as the Actor network, can be structured as a multilayer perceptron (MLP) or other common network structures in the related art. In this embodiment, the policy network is used to obtain state parameters as input and generate a control policy as output. The value network, also known as the Critic, can be structured as a convolutional neural network or other common network structures in the related art. In this embodiment, the value network is used to obtain the control policy generated by the policy network as input, evaluate and output the long-term cumulative reward corresponding to the control policy, and guide the policy network to optimize the previously output control policy based on the long-term cumulative reward. This process is repeated to obtain the optimal control policy.
[0108] In this embodiment, the above-mentioned deep reinforcement learning model can be obtained by firstly using the following offline training method:
[0109] Simulate the application scenarios of the target network system in a digital twin simulation environment, such as simulating industrial parks, smart hospitals, etc., to generate a number of simulation status data as samples. These simulation status data can be status data obtained by simulating sudden traffic, equipment failures, high-speed user movement, etc. in the simulated application scenarios.
[0110] Then, based on these simulated state data as samples, the pre-built initial model is trained offline using the proximal policy optimization (PPO) or deep deterministic policy gradient (DDPG) algorithm to obtain a trained deep learning reinforcement model. At this time, the deep learning reinforcement model can be deployed on the edge node so that the edge node can use the model to determine the target control strategy.
[0111] The specific training process based on sample data can be found in the relevant information on deep reinforcement learning models in the prior art and will not be described in detail here.
[0112] While the edge node is actually using the trained deep reinforcement learning model, the edge node can also perform online fine-tuning on the locally deployed deep reinforcement learning model. That is, the edge node can update the parameters of the locally deployed model based on real-time status data and the optimization effect indicators of the target network system after optimization according to the target control strategy, so that the locally deployed model can adapt to dynamic changes in the network (such as changes in the addition of new wireless access points or expansion of optical links).
[0113] According to the above deep reinforcement learning model, see Figure 3 In S102, the goal is to optimize the long-term cumulative reward corresponding to the control strategy. The method of determining the target control strategy according to the state parameters can be:
[0114] A1, processes state parameters according to the pre-built strategy network to obtain the current control strategy;
[0115] A2, processes the current control strategy according to the pre-built value network and obtains the current long-term cumulative reward corresponding to the current control strategy;
[0116] A3, if the current long-term cumulative reward does not meet the target optimization conditions, update the current control policy in the policy network based on the state parameters and the current long-term cumulative reward until the current long-term cumulative reward meets the target optimization conditions;
[0117] A4, determine the current control strategy whose current long-term cumulative reward meets the target optimization condition as the target control strategy.
[0118] Each control policy output by the policy network can include multiple executable adjustment actions for each terminal device and wireless access point connected to the edge node, as well as the corresponding action probability for each adjustment action. These adjustment actions can be discrete or continuous. Discrete actions can include maintaining the current connection, switching to the target access point, and triggering optical layer bandwidth reallocation. Other types can also be included. The target access point is a wireless access point different from the wireless access point currently connected to the terminal device. Continuous actions can include adjusting optical layer bandwidth, optical layer wavelength migration, adjusting signal power, adjusting signal strength, changing modulation parameters, and adjusting wireless channel allocation parameters.
[0119] In continuous operation, the range of optical layer wavelength migration can be between 0% and 100%, and the specific value can be specified in the control policy; adjusting the signal power is equivalent to adjusting the signal power of the wireless access point, and the adjustment range can be between -3dBm and +3dBm, and the specific value can be specified in the control policy.
[0120] Among them, discrete actions such as maintaining the current connection, switching to the target access point, triggering optical layer bandwidth reallocation, etc. can be targeted at each terminal device connected to the edge node. Figure 2 For example, the current control strategy may include maintaining the current connection action, switching to the target access point action, and triggering the optical layer bandwidth reallocation action for terminal device 1, and maintaining the current connection action, switching to the target access point action, and triggering the optical layer bandwidth reallocation action for terminal device 2.
[0121] The above continuous actions can be directed to each wireless access point connected to the edge node, or to the edge node. Figure 2 For example, the current control strategy may include an action of adjusting the signal strength and the signal power for wireless access point 1 , and may also include an action of adjusting the optical layer bandwidth for optical network unit 1 .
[0122] Whether it is a discrete action or a continuous action, each adjustment action can include the identification of the device targeted by the action, such as the identification of the targeted terminal device, the identification of the targeted wireless access point, and the identification of the targeted optical network unit, so as to perform the action on the corresponding device based on the identification.
[0123] In discrete actions, the action of switching to a target access point may further include an identifier corresponding to the target access point, so that the terminal device is switched to the specified target access point when the action is executed.
[0124] Each continuous action may further include an adjusted value. For example, an action of adjusting optical layer bandwidth may include the adjusted value of optical layer bandwidth, and an action of adjusting signal power may include the adjusted value of signal power.
[0125] It should be noted that the current control strategy may contain only discrete actions, or only continuous actions, or both discrete actions and continuous actions, and the discrete actions and continuous actions may be any combination of the multiple actions listed in the above examples.
[0126] In the method of this embodiment, the long-term cumulative reward can be determined based on the immediate reward value, which is obtained by processing multiple evaluation indicators according to a preset reward function.
[0127] Specifically, in A2, the value network predicts multiple evaluation metrics for the target network system after implementing the current control strategy. It then combines these metrics with a pre-defined reward function to determine the long-term cumulative reward corresponding to the current control strategy. Implementing the current control strategy can be understood as executing the adjustment action with the highest probability in the current control strategy for each device in the target network system (including terminal devices and wireless access points).
[0128] For example, among the three discrete actions of terminal device 1, the probability of maintaining the current connection action is the highest, so terminal device 1 remains connected to wireless access point 1. Among the three discrete actions of terminal device 2, the probability of switching to target access point 2 is the highest, so terminal device 2 is switched to connect to wireless access point 2.
[0129] The evaluation indicators used in this embodiment include but are not limited to switching success rate, end-to-end delay compliance rate, number of switching times, optical-wireless resource mismatch, etc., among which the optical-wireless resource mismatch can be determined based on the difference between the normalized optical resource utilization and the wireless resource utilization. The specific method can be referred to the relevant existing technology and will not be repeated here.
[0130] The reward function can specify the reward value corresponding to each evaluation indicator. As an example, it can be stipulated that when the handover success rate is greater than a certain threshold, a reward value of +10 is obtained, a reward value of +5 is obtained for each terminal device that meets the end-to-end delay compliance rate, a reward value corresponding to the number of handovers is -3 for each handover, and a reward value of -6 is obtained when the optical-wireless resource mismatch (also known as optical-wireless resource mismatch) is greater than a certain threshold.
[0131] As some examples, based on the above reward value distribution rules, the reward function can be expressed as follows, where R is the function value of the reward function calculated according to the above rules:
[0132] R = 0.4 × handover success rate + 0.3 × (1-optical and wireless resource mismatch) + 0.3 × end-to-end delay compliance rate.
[0133] Based on the set reward function, the value network can predict the values of various evaluation indicators of the target network system at different times in the future after the implementation of the current control strategy, and then use the reward function to calculate the evaluation indicators at different times to obtain the immediate reward values at different times in the future. The value network can discount and accumulate these immediate rewards over time according to a certain algorithm, and finally obtain the expected value of the cumulative reward of the current control strategy, which is the current long-term cumulative reward mentioned above.
[0134] In A3, the target optimization condition can be set as needed, for example, it can be set to that the current long-term cumulative reward is greater than or equal to a preset cumulative reward threshold, or it can be set to that the absolute value of the difference between the current long-term cumulative reward obtained this time and the current long-term cumulative reward obtained last time is less than or equal to a preset threshold.
[0135] If the current long-term cumulative reward obtained this time meets the target optimization conditions, then S102 is executed and the current control strategy output by the policy network most recently is output as the optimal target control strategy, and the process goes to step S103. If the current long-term cumulative reward obtained this time does not meet the target optimization conditions, then the current long-term cumulative reward obtained this time and the current control strategy obtained most recently can be input into the policy network, so that the policy network optimizes the current control strategy to obtain a new current control strategy, and so on, until the current long-term cumulative reward obtained at a certain time meets the target optimization conditions.
[0136] In some optional embodiments, the target control strategy may include multiple adjustment actions and an action probability corresponding to each adjustment action;
[0137] Optimizing the network resource allocation of the target network system according to the target control strategy in S103 may include:
[0138] Determining a target adjustment action in a target control strategy based on the corresponding action probability, and executing the target adjustment action to adjust the network resource allocation of the target network system;
[0139] The various adjustment actions include actions for adjusting the connection relationship between the terminal device and the target network system, and / or actions for adjusting parameters of the optical fiber network and the wireless network.
[0140] The action for adjusting the connection relationship between the terminal device and the target network system is equivalent to the discrete action of the aforementioned embodiment, and the action for adjusting the parameters of the optical fiber network and the wireless network is equivalent to the continuous action of the aforementioned embodiment.
[0141] As described in the above embodiment, in S103, after obtaining the target control strategy, the edge node can determine the adjustment action with the highest corresponding action probability in the target control strategy as the target adjustment action, and then execute the target adjustment action to adjust the network resource allocation of the target network system.
[0142] It should be noted that if the target control strategy includes multiple adjustment actions for multiple devices in the target network system, then when executing S103, for each device, the edge node needs to determine the target adjustment action for the device, and then execute the target adjustment action to change the connection relationship or parameters of the device.
[0143] The device herein may include the edge node executing the method of this embodiment, each wireless access point to which the edge node is connected, and each terminal device to which the edge node is connected. A terminal device connected to an edge node refers to a terminal device that accesses a wireless access point under the edge node via a wireless network.
[0144] by Figure 2 For example, Figure 2 Wireless access points 1 to 3 are wireless access points subordinate to (or connected to) optical network unit 1, and terminal devices 1 and 2 are terminal devices connected to optical network unit 1. The target control strategy determined by optical network unit 1 may include any one or more of a plurality of discrete actions for terminal device 1, a plurality of discrete actions for terminal device 2, a plurality of continuous actions for wireless access point 1, a plurality of continuous actions for wireless access point 2, a plurality of continuous actions for wireless access point 3, and a plurality of continuous actions for optical network unit 1. Furthermore, the target control strategy may also include a continuous action for terminal device 1 or a discrete action for wireless access point 1, without limitation to the above examples.
[0145] Based on the target control strategy, in S103, the optical network unit 1 as an edge node can find an action with the highest probability among multiple discrete actions corresponding to the terminal device 1 as the target adjustment action corresponding to the terminal device 1, such as executing a switching to the target access point action, and then executing the target adjustment action on the terminal device 1. It can also find an action with the highest probability among multiple continuous actions for the wireless access point 1 as the target adjustment action for the wireless access point 1, and then execute the target adjustment action on the wireless access point 1, such as executing a signal power adjustment action, and so on.
[0146] Optionally, performing a target adjustment action to adjust the network resource allocation of the target network system includes at least one of the following:
[0147] Adjustment method 1: controlling the first terminal device to switch from the first wireless access point to the second wireless access point, where the first wireless access point is the wireless access point currently accessed by the first terminal device;
[0148] Adjustment method 2: Migrating part of the network services of the first optical network unit connected to the third wireless access point on the current wavelength channel to another idle wavelength channel, where the third wireless access point and the second wireless access point are the same or different wireless access points;
[0149] Adjustment method three: adjusting the optical layer bandwidth resources allocated to each connected wireless access point by the second optical network unit, where the second optical network unit and the first optical network unit are the same or different optical network units.
[0150] Adjustment method one is equivalent to executing a switching action to a target access point on the first terminal device, wherein the first wireless access point is the wireless access point that the first terminal device originally connected to before executing the action, and the second wireless access point is equivalent to the wireless access point to be switched to specified in the switching action to the target access point.
[0151] The first terminal device may be any terminal device connected to the edge node that executes the optimization method of this embodiment. Figure 2 In the example, the first terminal device may be terminal device 2, the first wireless access point may be wireless access point 1 therein, and the second wireless access point is equivalent to wireless access point 2 therein. By executing the target access point action on terminal device 2, terminal device 2 may be changed to access wireless access point 2.
[0152] When controlling the first terminal device to switch to the second wireless access point according to adjustment mode 1, the edge node executing the method of this embodiment may send control signaling for switching to the second wireless access point to the first terminal device through the first wireless access point. The first terminal device may then respond to the control signaling by sending a handshake signal to the second wireless access point. The handshake signal may be a Request To Send / Clear To Send (RTS / CTS) handshake signal to trigger the second wireless access point to synchronously update its own terminal association table. After the second wireless access point has updated its own terminal association table, it indicates that the first terminal device has completed the handover from the first wireless access point to the second wireless access point.
[0153] Each wireless access point may have a terminal association table, which may be used to record the device identifications of terminal devices accessing the wireless access point.
[0154] Optionally, in each handover to a target access point action, the specific target access point may be determined by the policy network through signal strength, channel quality, load balancing, and the like.
[0155] The first wireless access point and the second wireless access point may belong to the same optical network unit or to different optical network units.
[0156] Adjustment method 2 is equivalent to triggering optical layer bandwidth reallocation on the edge node (ie, the optical network unit executing the method of this embodiment).
[0157] The network service migration in adjustment method 2 can be specifically performed by the first optical network unit. The first optical network unit can determine an idle wavelength channel from its multiple wavelength channels and a current wavelength channel carrying the largest traffic. Then, the first optical network unit can migrate a portion of the network services carried by the current wavelength channel to the idle wavelength channel according to a certain ratio. The specific migration method can be found in the relevant existing technology and is not described in detail here.
[0158] In some embodiments, the third wireless access point and the second wireless access point may be two different wireless access points. That is, when executing S103, the edge node executing the method may independently switch the wireless access point to which the first terminal device is connected, and reallocate the optical layer bandwidth of the first optical network unit.
[0159] In some embodiments, the third wireless access point and the second wireless access point may be the same wireless access point. After the edge node switches the first terminal device from the first wireless access point to the second wireless access point, if it is found that the load of the optical network unit connected to the second wireless access point is too high after the switch, the optical network unit connected to the second wireless access point may be used as the first optical network unit, and optical layer bandwidth reallocation may be performed on the first optical network unit according to adjustment method two.
[0160] Adjustment method three is equivalent to performing an action of adjusting the optical layer bandwidth on the second optical network unit.
[0161] The optical layer bandwidth adjustment action may include the identifiers of one or more wireless access points under the second optical network unit whose optical layer bandwidth needs to be adjusted, as well as the target optical layer bandwidth values of these wireless access points. When performing adjustment according to adjustment method three, the second optical network unit can, based on the optical layer bandwidth adjustment action, identify the wireless access points that need to be adjusted among its connected wireless access points and then adjust the optical layer bandwidth allocated to these wireless access points to the corresponding target values.
[0162] In some embodiments, the second optical network unit and the first optical network unit may be two different optical network units, that is, when executing S103, the optical layer bandwidth reallocation action and the optical layer bandwidth adjustment action may be performed on the two different optical network units respectively.
[0163] In some embodiments, the second optical network unit and the first optical network unit may be the same optical network unit. That is, when executing S103, on the one hand, part of the network services of the first optical network unit may be migrated from the current wavelength channel to the idle wavelength channel, and on the other hand, the optical layer bandwidth resources allocated by the first optical network unit to different wireless access points may be adjusted to meet the traffic requirements of different wireless access points of the first optical network unit.
[0164] In some embodiments, after executing corresponding adjustment actions according to the target control strategy, the edge node may also execute other actions in conjunction with the transformation of the network system after executing the action to achieve collaborative optimization:
[0165] Combine Figure 2In the example, according to the target control strategy, the terminal device 1 is switched to the wireless access point 2, and the terminal device 1 is switched from connecting to the wireless access point 1 to connecting to the wireless access point 2. After the switch is completed, if the corresponding edge node (i.e., optical network unit 1) finds that its own λ1 wavelength utilization rate is greater than a certain threshold, for example, greater than 90%, the edge node can trigger the execution of the following optical layer linkage adjustment action:
[0166] First, it sends a wavelength migration instruction to the OLT to which it is connected, migrating 20% of the low-priority services (such as surveillance video services) in wireless access point 2 to another idle wavelength, such as wavelength λ5.
[0167] Second, the signal transmission power of the wireless access point 2 is synchronously adjusted, for example, the signal transmission power of the wireless access point 2 is increased by 2 dBm to compensate for the signal fluctuation during the wavelength migration.
[0168] In some optional embodiments, the following steps may be performed after each execution of the target control strategy:
[0169] Obtain optimization effect indicators;
[0170] Adjust the reward function based on the optimization effect index;
[0171] The optimization effect indicator includes at least one of the switching time and the service interruption duration.
[0172] The switching time here may be the time taken for the terminal device to switch from the current wireless access point to the target access point, and the service interruption duration may refer to the duration of the service interruption phenomenon occurring during the switching process.
[0173] Taking the optimization effect indicators including switching time and service interruption duration as an example, in the above embodiment, the edge node can use the aforementioned value network to predict the switching time prediction value and service interruption duration prediction value of each switched terminal device after executing the target control strategy. On the other hand, after executing the target control strategy, the actual value of the optimization effect indicator of the switched terminal device is counted, that is, the actual switching time and service interruption duration of these terminal devices, and then the actual value is compared with the predicted value. If the actual value of a terminal device is less than or equal to the predicted value, it is determined that the terminal device meets the standard. If the actual value of a terminal device is greater than the predicted value, it is determined that the terminal device does not meet the standard.
[0174] Finally, if the proportion of substandard terminal devices among the switched terminal devices is greater than a certain threshold, the reward function can be adjusted; if the proportion of substandard terminal devices is less than or equal to the threshold, the reward function may not be adjusted.
[0175] The reward function can be adjusted by adding new evaluation indicators to the reward function or changing the scores corresponding to the evaluation indicators in the original reward function. The specific adjustment method and magnitude can be determined by the relevant users and will not be elaborated here.
[0176] Optionally, the model used to determine the target control strategy can be stress-tested regularly to simulate extreme scenarios (such as simulating a scenario where 100 users initiate a video conference at the same time) to verify the robustness of the model.
[0177] A second aspect of the present application provides a network system including a wireless network and an optical fiber network;
[0178] The wireless network includes terminal devices and wireless access points for accessing the terminal devices;
[0179] The optical fiber network includes an optical line terminal and an optical network unit connected by optical fibers, and the optical network unit and the wireless access point are connected by optical fibers;
[0180] Optical Network Units are used for:
[0181] Obtaining state parameters of a target network system composed of a fiber optic network and a wireless network, the state parameters including wireless parameters reflecting a network environment of the wireless network, optical layer parameters reflecting a network environment of the fiber optic network, and terminal parameters reflecting terminal requirements within a future preset time period;
[0182] Taking the optimization of the long-term cumulative reward corresponding to the control strategy as the goal, a target control strategy is determined according to the state parameters. The long-term cumulative reward is determined by multiple evaluation indicators, and the multiple evaluation indicators include any number of a first indicator reflecting the success rate of the control strategy, a second indicator reflecting the resources consumed by the control strategy, and a third indicator reflecting the effect achieved by the control strategy;
[0183] Optimize the network resource allocation of the target network system according to the target control strategy.
[0184] Optionally, the optical network unit is further used for at least one of the following:
[0185] In response to a predetermined control cycle, executing the step of obtaining a state parameter of a target network system consisting of an optical fiber network and a wireless network;
[0186] In response to detecting a target trigger event, executing the step of obtaining a state parameter of a target network system consisting of an optical fiber network and a wireless network;
[0187] The target trigger event includes an optical layer parameter satisfying a first trigger condition and / or a wireless parameter satisfying a second trigger condition.
[0188] Optionally, the optical network unit is further configured to:
[0189] The fault risk information of the optical fiber network is determined according to the optical layer parameters. The fault risk information at least represents the probability of a potential fault existing in the optical fiber network.
[0190] Optionally, the target control strategy includes multiple adjustment actions and an action probability corresponding to each adjustment action;
[0191] The optical network unit optimizes the network resource allocation of the target network system according to the target control strategy, including:
[0192] Determining a target adjustment action in a target control strategy based on the corresponding action probability, and executing the target adjustment action to adjust the network resource allocation of the target network system;
[0193] The various adjustment actions include actions for adjusting the connection relationship between the terminal device and the target network system, and / or actions for adjusting parameters of the optical fiber network and the wireless network.
[0194] Optionally, the optical network unit performs a target adjustment action to adjust the network resource allocation of the target network system, including at least one of the following:
[0195] Controlling the first terminal device to switch from the first wireless access point to the second wireless access point, where the first wireless access point is the wireless access point currently accessed by the first terminal device;
[0196] Migrating part of the network services of the first optical network unit connected to the third wireless access point on the current wavelength channel to another idle wavelength channel, where the third wireless access point and the second wireless access point are the same or different wireless access points;
[0197] The optical layer bandwidth resources allocated to each connected wireless access point by the second optical network unit are adjusted, and the second optical network unit and the first optical network unit are the same or different optical network units.
[0198] Optionally, long-term cumulative rewards are determined by multiple evaluation indicators, including:
[0199] Long-term cumulative rewards are determined based on the immediate reward value, which is obtained by processing multiple evaluation indicators according to the preset reward function;
[0200] After optimizing the network resource allocation of the target network system according to the target control strategy, the optical network unit is also used to:
[0201] Obtain optimization effect indicators;
[0202] Adjust the reward function based on the optimization effect index;
[0203] The optimization effect indicator includes at least one of the switching time and the service interruption duration.
[0204] Optionally, the optical network unit determines a target control strategy based on the state parameters with the goal of optimizing the long-term cumulative reward corresponding to the control strategy, including:
[0205] Process state parameters according to the pre-built strategy network to obtain the current control strategy;
[0206] Process the current control strategy according to the pre-built value network to obtain the current long-term cumulative reward corresponding to the current control strategy;
[0207] If the current long-term cumulative reward does not meet the target optimization conditions, the current control policy is updated in the policy network based on the state parameters and the current long-term cumulative reward until a current control policy is obtained in which the current long-term cumulative reward meets the target optimization conditions;
[0208] The current control strategy whose current long-term cumulative reward meets the target optimization condition is determined as the target control strategy.
[0209] Optionally, the optical layer parameter includes at least one of optical fiber link delay, optical power fluctuation, wavelength occupancy, bandwidth utilization, and bit error rate;
[0210] And / or, the wireless parameter includes at least one of a channel signal-to-noise ratio, co-channel interference intensity, and the number of user accesses.
[0211] Optionally, the terminal parameters include at least one of an expected movement trajectory parameter and a service traffic type of the terminal device, and the expected movement trajectory parameter is determined based on a historical location parameter, a movement speed parameter, and a movement acceleration parameter of the terminal device.
[0212] The structure of the above target network system can be found in Figure 2 The specific working principle can be found in the relevant steps of the optimization method in the above embodiment and will not be described in detail.
[0213] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referenced to each other.
[0214] For the convenience of description, the above systems or devices are described as being divided into various modules or units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0215] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.
[0216] Finally, it should be noted that, in this document, relational terms such as first, second, third, and fourth are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0217] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A network optimization method, characterized in that: include: Obtaining state parameters of a target network system consisting of a fiber optic network and a wireless network, the state parameters including wireless parameters reflecting a network environment of the wireless network, optical layer parameters reflecting a network environment of the fiber optic network, and terminal parameters reflecting requirements of terminal devices within a future preset time period; With the goal of optimizing the long-term cumulative reward corresponding to the control strategy, a target control strategy is determined according to the state parameter, wherein the long-term cumulative reward is determined by a plurality of evaluation indicators, wherein the plurality of evaluation indicators include any multiple of a first indicator reflecting the success rate of the control strategy, a second indicator reflecting the resources consumed by the control strategy, and a third indicator reflecting the effect achieved by the control strategy; The network resource allocation of the target network system is optimized according to the target control strategy.
2. The method according to claim 1, characterized in that Also includes at least one of the following: In response to a preset control cycle, executing the step of obtaining the state parameters of the target network system consisting of the optical fiber network and the wireless network; In response to detecting a target triggering event, executing the step of obtaining the state parameters of the target network system consisting of the optical fiber network and the wireless network; The target trigger event includes an optical layer parameter satisfying a first trigger condition and / or a wireless parameter satisfying a second trigger condition.
3. The method according to claim 1, characterized in that Also includes: Fault risk information of the optical fiber network is determined according to the optical layer parameters, where the fault risk information at least represents a probability of a potential fault existing in the optical fiber network.
4. The method according to claim 1, wherein The target control strategy includes multiple adjustment actions and an action probability corresponding to each adjustment action; Optimizing the network resource allocation of the target network system according to the target control strategy includes: determining a target adjustment action in the target control strategy based on the corresponding action probability, and executing the target adjustment action to adjust the network resource allocation of the target network system; The multiple adjustment actions include actions for adjusting the connection relationship between the terminal device and the target network system, and / or actions for adjusting bandwidth resources of the optical fiber network and the wireless network.
5. The method according to claim 4, characterized in that The performing of the target adjustment action to adjust the network resource allocation of the target network system includes at least one of the following: Controlling a first terminal device to switch from a first wireless access point to a second wireless access point, where the first wireless access point is the wireless access point currently accessed by the first terminal device; Migrating part of the network services of the first optical network unit connected to the third wireless access point on the current wavelength channel to another idle wavelength channel, wherein the third wireless access point and the second wireless access point are the same or different wireless access points; Adjust the optical layer bandwidth resources allocated to each connected wireless access point by the second optical network unit, where the second optical network unit and the first optical network unit are the same or different optical network units.
6. The method according to claim 1, characterized in that The long-term cumulative reward is determined by multiple evaluation indicators, including: The long-term cumulative reward is determined based on the immediate reward value, which is obtained by processing multiple evaluation indicators according to a preset reward function; After optimizing the network resource allocation of the target network system according to the target control strategy, the method further includes: Obtain optimization effect indicators; Adjusting the reward function according to the optimization effect indicator; The optimization effect indicator includes at least one of switching time and service interruption duration.
7. The method according to claim 1, characterized in that The method of determining a target control strategy based on the state parameters with the goal of optimizing the long-term cumulative reward corresponding to the control strategy includes: Processing the state parameters according to a pre-built strategy network to obtain a current control strategy; Processing the current control strategy according to a pre-built value network to obtain a current long-term cumulative reward corresponding to the current control strategy; If the current long-term cumulative reward does not satisfy the target optimization condition, updating the current control strategy in the strategy network based on the state parameter and the current long-term cumulative reward until a current control strategy in which the current long-term cumulative reward satisfies the target optimization condition is obtained; The current control strategy whose current long-term cumulative reward satisfies the target optimization condition is determined as the target control strategy.
8. The method according to claim 1, characterized in that The optical layer parameters include at least one of optical fiber link delay, optical power fluctuation, wavelength occupancy, available wavelength resources, optical fiber link delay standard deviation, bandwidth utilization and bit error rate; And / or, the wireless parameter includes at least one of a channel signal-to-noise ratio, co-channel interference intensity, number of user accesses, and wireless access point load rate.
9. The method according to claim 1, characterized in that The terminal parameters include at least one of an expected movement trajectory parameter of the terminal device and a service traffic type, and the expected movement trajectory parameter is determined according to a historical location parameter, a movement speed parameter, and a movement acceleration parameter of the terminal device.
10. A network system, characterized in that: including wireless and fiber optic networks; The wireless network includes a terminal device and a wireless access point for accessing the terminal device; The optical fiber network includes an optical line terminal and an optical network unit connected by optical fibers, and the optical network unit and the wireless access point are connected by optical fibers; The optical network unit is used for: Obtaining state parameters of a target network system consisting of a fiber optic network and a wireless network, the state parameters including wireless parameters reflecting a network environment of the wireless network, optical layer parameters reflecting a network environment of the fiber optic network, and terminal parameters reflecting terminal demand within a future preset time period; With the goal of optimizing the long-term cumulative reward corresponding to the control strategy, a target control strategy is determined according to the state parameter, wherein the long-term cumulative reward is determined by a plurality of evaluation indicators, wherein the plurality of evaluation indicators include any multiple of a first indicator reflecting the success rate of the control strategy, a second indicator reflecting the resources consumed by the control strategy, and a third indicator reflecting the effect achieved by the control strategy; The network resource allocation of the target network system is optimized according to the target control strategy.